Identifying bot activity in GitHub pull request and issue comments

by   Mehdi Golzadeh, et al.

Development bots are used on Github to automate repetitive activities. Such bots communicate with human actors via issue comments and pull request comments. Identifying such bot comments allows preventing bias in socio-technical studies related to software development. To automate their identification, we propose a classification model based on natural language processing. Starting from a balanced ground-truth dataset of 19,282 PR and issue comments, we encode the comments as vectors using a combination of the bag of words and TF-IDF techniques. We train a range of binary classifiers to predict the type of comment (human or bot) based on this vector representation. A multinomial Naive Bayes classifier provides the best results. Its performance on a test set containing 50 and F1 score of 0.88. Although the model shows a promising result on the pull request and issue comments, further work is required to generalize the model on other types of activities, like commit messages and code reviews.



There are no comments yet.


page 1

page 2

page 3

page 4


A ground-truth dataset and classification model for detecting bots in GitHub issue and PR comments

Bots are frequently used in Github repositories to automate repetitive a...

Evaluating a bot detection model on git commit messages

Detecting the presence of bots in distributed software development activ...

Descriptions of issues and comments for predicting issue success in software projects

Software development tasks must be performed successfully to achieve sof...

A Transfer Learning Approach for Dialogue Act Classification of GitHub Issue Comments

Social coding platforms, such as GitHub, serve as laboratories for study...

Reading Between the Demographic Lines: Resolving Sources of Bias in Toxicity Classifiers

The censorship of toxic comments is often left to the judgment of imperf...

Automating the Removal of Obsolete TODO Comments

TODO comments are very widely used by software developers to describe th...

Pull Request Latency Explained: An Empirical Overview

Pull request latency evaluation is an essential application of effort ev...
This week in AI

Get the week's most popular data science and artificial intelligence research sent straight to your inbox every Saturday.