Query Expansion for Patent Searching using Word Embedding and Professional Crowdsourcing

by   Arthi Krishna, et al.

The patent examination process includes a search of previous work to verify that a patent application describes a novel invention. Patent examiners primarily use keyword-based searches to uncover prior art. A critical part of keyword searching is query expansion, which is the process of including alternate terms such as synonyms and other related words, since the same concepts are often described differently in the literature. Patent terminology is often domain specific. By curating technology-specific corpora and training word embedding models based on these corpora, we are able to automatically identify the most relevant expansions of a given word or phrase. We compare the performance of several automated query expansion techniques against expert specified expansions. Furthermore, we explore a novel mechanism to extract related terms not just based on one input term but several terms in conjunction by computing their centroid and identifying the nearest neighbors to this centroid. Highly skilled patent examiners are often the best and most reliable source of identifying related terms. By designing a user interface that allows examiners to interact with the word embedding suggestions, we are able to use these interactions to power crowdsourced modes of related terms. Learning from users allows us to overcome several challenges such as identifying words that are bleeding edge and have not been published in the corpus yet. This paper studies the effectiveness of word embedding and crowdsourced models across 11 disparate technical areas.


Merchandise Recommendation for Retail Events with Word Embedding Weighted Tf-idf and Dynamic Query Expansion

To recommend relevant merchandises for seasonal retail events, we rely o...

Relevance-based Word Embedding

Learning a high-dimensional dense representation for vocabulary terms, a...

Detecting New Word Meanings: A Comparison of Word Embedding Models in Spanish

Semantic neologisms (SN) are defined as words that acquire a new word me...

Vocab-Expander: A System for Creating Domain-Specific Vocabularies Based on Word Embeddings

In this paper, we propose Vocab-Expander at https://vocab-expander.com, ...

LEXpander: applying colexification networks to automated lexicon expansion

Recent approaches to text analysis from social media and other corpora r...

Learning Word Relatedness over Time

Search systems are often focused on providing relevant results for the "...

Quality-aware skill translation models for expert finding on StackOverflow

StackOverflow has become an emerging resource for talent recognition in ...

Please sign up or login with your details

Forgot password? Click here to reset