Microsoft Corporation
DOCUMENT BODY VECTORIZATION AND NOISE-CONTRASTIVE TRAINING

Last updated: 22 Jun 2022

Abstract:

Document embedding vectors for each document of a corpus may be generated by combining embedding vectors for document subparts, thereby yielding a final embedding vector for the document. A machine learning model is trained using a query corpus and the document corpus, where the model generates a ranking score for a given (query, document) pair. During training, rankings scores are generated using the model, such that the training dataset is further refined using the generated ranking scores. For example, top documents and a negative document may be determined for a given query and subsequently used as training data. Multiple negative documents may therefore be determined for a given query. A negative document for a given query may be determined from the negative documents using noise-contrastive estimation. Such determined negative documents may be evaluated using a loss function during model training, thereby yielding a more robust model for search processing.

Status:

Application

Type:

Utility

Filling date:

19 Mar 2021

Issue date:

9 Jun 2022

Full patent description

Patent application document

Microsoft Corporation DOCUMENT BODY VECTORIZATION AND NOISE-CONTRASTIVE TRAINING

Abstract:

Microsoft Corporation
DOCUMENT BODY VECTORIZATION AND NOISE-CONTRASTIVE TRAINING