Inventors
Luciano De Andrade Barbosa, Srinivas Bangalore
Publication date
2015/7/14
Patent office
US
Patent number
9081760
Application number
13042890
Description
Disclosed herein are systems, methods, and non-transitory computer-readable storage media for collecting web data in order to create diverse language models. A system configured to practice the method first crawls, such as via a crawler operating on a computing device, a set of documents in a network of interconnected devices according to a visitation policy, wherein the visitation policy is configured to focus on novelty regions for a current language model built from previous crawling cycles by crawling documents whose vocabulary considered likely to fill gaps in the current language model. A language model from a previous cycle can be used to guide the creation of a language model in the following cycle. The novelty regions can include documents with high perplexity values over the current language model.
Total citations
2014201520162017201820192020202120222023202413151521464538504718
Scholar articles