Generation-Focused Table-Based Intermediate Pre-training for Free-Form Question Answering

Shi, Peng; Ng, Patrick; Feng, Nan; Zhu, Henghui; Wang, Jun; Jiang, Jiarong; Li, Alexander Hanbo; Chakravarti, Rishav; Weidner, Donald J.; Xiang, Bing; Wang, Zhiguo

doi:10.1609/aaai.v36i10.21382

Cited by 3 publications

(2 citation statements)

References 37 publications

Supporting

Mentioning

Contrasting

Order By: Relevance

“…Our pretraining objectives include low-level tasks that are more focused on retrieving the underlying data from the chart images and high-level tasks that align closely with the downstream tasks. (Shi et al, 2022) to generate synthetic open-ended QA pairs. Specifically, a T5 model (Raffel et al, 2020) pretrained on SQuAD (Rajpurkar et al, 2016) is employed to generate an open-ended question for each summary.…”

Section: Pretraining Objectivesmentioning

confidence: 99%

UniChart: A Universal Vision-language Pretrained Model for Chart Comprehension and Reasoning

Masry,

Kavehzadeh,

Long

et al. 2023

Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing

View full text Add to dashboard Cite

Charts are widely used for data analysis, providing visual representations and insights into complex data. To facilitate chart-based data analysis using natural language, several downstream tasks have been introduced recently such as chart question answering and chart summarization. However, existing methods for these tasks often rely on pretraining on language or vision-language tasks, neglecting the explicit modeling of chart structures (e.g., how chart elements are related to each other). To address this, we first build a large corpus of charts covering diverse topics and visual styles. We then present UniChart, a pretrained model for chart comprehension and reasoning. UniChart encodes the relevant text, data, and visual elements of charts and then uses a chart-grounded text decoder for text generation. We propose several chart-specific pretraining tasks that include: (i) low-level tasks to extract the visual elements (e.g., bars, lines) and data from charts, and (ii) high-level tasks to acquire chart understanding and reasoning skills. Our experiments demonstrate that pretraining UniChart on a large corpus with chart-specific objectives, followed by fine-tuning, yields state-of-the-art performance on four downstream tasks. Moreover, our model exhibits superior generalizability to unseen chart corpus, surpassing previous approaches that lack chart-specific objectives and utilize limited chart resources.

show abstract

Section: Pretraining Objectivesmentioning

confidence: 99%

UniChart: A Universal Vision-language Pretrained Model for Chart Comprehension and Reasoning

Masry,

Kavehzadeh,

Long

et al. 2023

Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing

View full text Add to dashboard Cite

show abstract

“…pretrains the PLMs with structure-augmented objectives (Herzig et al 2020;Deng et al 2021). Specifically, these works usually design unsupervised or weakly-supervised objectives for implicitly modeling the database structures with external or synthetic data corpus (Yu et al 2021b;Shi et al 2022). Although effective, further pretraining a large PLM can incur substantial costs and extra overheads (Yu et al 2021a).…”

Section: Introductionmentioning

confidence: 99%

Unveiling the Black Box of PLMs with Semantic Anchors: Towards Interpretable Neural Semantic Parsing

Nie

Jiuding

Wang

et al. 2023

AAAI

View full text Add to dashboard Cite

The recent prevalence of pretrained language models (PLMs) has dramatically shifted the paradigm of semantic parsing, where the mapping from natural language utterances to structured logical forms is now formulated as a Seq2Seq task. Despite the promising performance, previous PLM-based approaches often suffer from hallucination problems due to their negligence of the structural information contained in the sentence, which essentially constitutes the key semantics of the logical forms. Furthermore, most works treat PLM as a black box in which the generation process of the target logical form is hidden beneath the decoder modules, which greatly hinders the model's intrinsic interpretability. To address these two issues, we propose to incorporate the current PLMs with a hierarchical decoder network. By taking the first-principle structures as the semantic anchors, we propose two novel intermediate supervision tasks, namely Semantic Anchor Extraction and Semantic Anchor Alignment, for training the hierarchical decoders and probing the model intermediate representations in a self-adaptive manner alongside the fine-tuning process. We conduct intensive experiments on several semantic parsing benchmarks and demonstrate that our approach can consistently outperform the baselines. More importantly, by analyzing the intermediate representations of the hierarchical decoders, our approach also makes a huge step toward the interpretability of PLMs in the domain of semantic parsing.

show abstract

Around the GLOBE: Numerical Aggregation Question-answering on Heterogeneous Genealogical Knowledge Graphs with Deep Neural Networks

Suissa

Zhitomirsky‐Geffet

Elmalech

2023

J. Comput. Cult. Herit.

View full text Add to dashboard Cite

One of the key AI tools for textual corpora exploration is natural language question-answering (QA). Unlike keyword-based search engines, QA algorithms receive and process natural language questions and produce precise answers to these questions, rather than long lists of documents that need to be manually scanned by the users. State-of-the-art QA algorithms based on DNNs were successfully employed in various domains. However, QA in the genealogical domain is still underexplored, while researchers in this field (and other fields in humanities and social sciences) can highly benefit from the ability to ask questions in natural language, receive concrete answers and gain insights hidden within large corpora. While some research has been recently conducted for factual QA in the genealogical domain, to the best of our knowledge, there is no previous research on the more challenging task of numerical aggregation QA (i.e., answering questions combining aggregation functions, e.g., count, average, max). Numerical aggregation QA is critical for distant reading and analysis for researchers (and the general public) interested in investigating cultural heritage domains. Therefore, in this study, we present a new end-to-end methodology for numerical aggregation QA for genealogical trees that includes: 1) an automatic method for training dataset generation; 2) a transformer-based table selection method, and 3) an optimized transformer-based numerical aggregation QA model. The findings indicate that the proposed architecture, GLOBE, outperforms the state-of-the-art models and pipelines by achieving 87% accuracy for this task compared to only 21% by current state-of-the-art models. This study may have practical implications for genealogical information centers and museums, making genealogical data research easy and scalable for experts as well as the general public.

show abstract

Generation-Focused Table-Based Intermediate Pre-training for Free-Form Question Answering

Cited by 3 publications

References 37 publications

UniChart: A Universal Vision-language Pretrained Model for Chart Comprehension and Reasoning

UniChart: A Universal Vision-language Pretrained Model for Chart Comprehension and Reasoning

Unveiling the Black Box of PLMs with Semantic Anchors: Towards Interpretable Neural Semantic Parsing

Around the GLOBE: Numerical Aggregation Question-answering on Heterogeneous Genealogical Knowledge Graphs with Deep Neural Networks

Contact Info

Product

Resources

About