Publications

DynamicToC: Persona-based Table of Contents for Consumption of Long Documents

North American Chapter of the Association for Computational Linguistics (NAACL)

Publication date: July 15, 2022

Himanshu Maheshwari, Nethraa Shivakumar, Shelly Jain, Tanvi Karandikar, Navita Goyal, Vinay Aggarwal, Sumit Shekhar

Long documents like contracts, financial documents, etc., are often tedious to read through. Linearly consuming (via scrolling or navigation through default table of content) these documents is time-consuming and challenging. These documents are also authored to be consumed by varied entities (referred to as persona in the paper) interested in only certain parts of the document. In this work, we describe DYNAMICTOC, a dynamic table of content-based navigator, to aid in the task of non-linear, persona-based document consumption. DYNAMICTOC highlights sections of interest in the document as per the aspects relevant to different personas. DYNAMICTOC is augmented with short questions to assist the users in understanding underlying content. This uses a novel deep-reinforcement learning technique to generate questions on these persona-clustered paragraphs. Human and automatic evaluations suggest the efficacy of both end-to-end pipeline and different components of DYNAMICTOC.

Learn More