Time-sync Video Tag Extraction Using Semantic Association Graph

by   Wenmian Yang, et al.

Time-sync comments reveal a new way of extracting the online video tags. However, such time-sync comments have lots of noises due to users' diverse comments, introducing great challenges for accurate and fast video tag extractions. In this paper, we propose an unsupervised video tag extraction algorithm named Semantic Weight-Inverse Document Frequency (SW-IDF). Specifically, we first generate corresponding semantic association graph (SAG) using semantic similarities and timestamps of the time-sync comments. Second, we propose two graph cluster algorithms, i.e., dialogue-based algorithm and topic center-based algorithm, to deal with the videos with different density of comments. Third, we design a graph iteration algorithm to assign the weight to each comment based on the degrees of the clustered subgraphs, which can differentiate the meaningful comments from the noises. Finally, we gain the weight of each word by combining Semantic Weight (SW) and Inverse Document Frequency (IDF). In this way, the video tags are extracted automatically in an unsupervised way. Extensive experiments have shown that SW-IDF (dialogue-based algorithm) achieves 0.4210 F1-score and 0.4932 MAP (Mean Average Precision) in high-density comments, 0.4267 F1-score and 0.3623 MAP in low-density comments; while SW-IDF (topic center-based algorithm) achieves 0.4444 F1-score and 0.5122 MAP in high-density comments, 0.4207 F1-score and 0.3522 MAP in low-density comments. It has a better performance than the state-of-the-art unsupervised algorithms in both F1-score and MAP.


page 7

page 21


Interactive Variance Attention based Online Spoiler Detection for Time-Sync Comments

Nowadays, time-sync comment (TSC), a new form of interactive comments, h...

Predicting Different Types of Subtle Toxicity in Unhealthy Online Conversations

This paper investigates the use of machine learning models for the class...

About Evaluation of F1 Score for RECENT Relation Extraction System

This document contains a discussion of the F1 score evaluation used in t...

Optimize_Prime@DravidianLangTech-ACL2022: Abusive Comment Detection in Tamil

This paper tries to address the problem of abusive comment detection in ...

STACC: Code Comment Classification using SentenceTransformers

Code comments are a key resource for information about software artefact...

Automatically Detecting Cyberbullying Comments on Online Game Forums

Online game forums are popular to most of game players. They use it to c...

Atypical lexical abbreviations identification in Russian medical texts

Abbreviation is a method of word formation that aims to construct the sh...

Please sign up or login with your details

Forgot password? Click here to reset