International Business Machines Corporation
DETECTING AND PROCESSING SECTIONS SPANNING PROCESSED DOCUMENT PARTITIONS

Last updated:

Abstract:

Aspects of the invention include detecting and processing sections spanning processed document partitions by caching a document partition. The document partition includes metadata indicating that the document partition is a portion of a whole document. Aspects also include pairing a candidate paragraph from the document partition with a cached paragraph segment and determining, using a coherence model, a probability that the candidate paragraph and the cached paragraph segment constitute a semantically coherent paragraph. Aspects further include discarding the cached paragraph segment and processing the candidate paragraph and the cached paragraph segment separately based on a determination that the probability is less than a threshold level and processing the candidate paragraph and the cached paragraph segment together as a cross-partition paragraph based on a determination that the probability is greater than the threshold level.

Status:
Application
Type:

Utility

Filling date:

27 Jul 2020

Issue date:

27 Jan 2022