LiveChat: A Large-Scale Personalized Dialogue Dataset Automatically Constructed from Live Streaming

06/14/2023
by   Jingsheng Gao, et al.
0

Open-domain dialogue systems have made promising progress in recent years. While the state-of-the-art dialogue agents are built upon large-scale text-based social media data and large pre-trained models, there is no guarantee these agents could also perform well in fast-growing scenarios, such as live streaming, due to the bounded transferability of pre-trained models and biased distributions of public datasets from Reddit and Weibo, etc. To improve the essential capability of responding and establish a benchmark in the live open-domain scenario, we introduce the LiveChat dataset, composed of 1.33 million real-life Chinese dialogues with almost 3800 average sessions across 351 personas and fine-grained profiles for each persona. LiveChat is automatically constructed by processing numerous live videos on the Internet and naturally falls within the scope of multi-party conversations, where the issues of Who says What to Whom should be considered. Therefore, we target two critical tasks of response modeling and addressee recognition and propose retrieval-based baselines grounded on advanced techniques. Experimental results have validated the positive effects of leveraging persona profiles and larger average sessions per persona. In addition, we also benchmark the transferability of advanced generation-based models on LiveChat and pose some future directions for current challenges.

READ FULL TEXT

page 5

page 15

research
03/08/2022

Towards Building an Open-Domain Dialogue System Incorporated with Internet Memes

In recent years, Internet memes have been widely used in online chatting...
research
04/24/2023

SocialDial: A Benchmark for Socially-Aware Dialogue Systems

Dialogue systems have been widely applied in many scenarios and are now ...
research
09/20/2021

PLATO-XL: Exploring the Large-scale Pre-training of Dialogue Generation

To explore the limit of dialogue generation pre-training, we present the...
research
05/22/2023

Towards Robust Personalized Dialogue Generation via Order-Insensitive Representation Regularization

Generating persona consistent dialogue response is important for develop...
research
08/27/2022

MDIA: A Benchmark for Multilingual Dialogue Generation in 46 Languages

Owing to the lack of corpora for low-resource languages, current works o...
research
02/17/2021

Integrating Pre-trained Model into Rule-based Dialogue Management

Rule-based dialogue management is still the most popular solution for in...
research
04/25/2023

OFAR: A Multimodal Evidence Retrieval Framework for Illegal Live-streaming Identification

Illegal live-streaming identification, which aims to help live-streaming...

Please sign up or login with your details

Forgot password? Click here to reset