Abstract:This paper develops the first question answering dataset (DrugEHRQA) containing question-answer pairs from both structured tables and unstructured notes from a publicly available Electronic Health Record (EHR). EHRs contain patient records, stored in structured tables and unstructured clinical notes. The information in structured and unstructured EHRs is not strictly disjoint: information may be duplicated, contradictory, or provide additional context between these sources. Our dataset has medication-related q… Show more
“…Therefore, there is a need for evaluating the quality of the user search experience when searching for information. And this is especially necessary because of the emergence and popularity of new search techniques such as Elastic search [16] [26], Question Answering over free text [3] [6] [4] [2], and Question Answering over knowledge graphs [8] [9] [27] [30]. The search techniques are one-field one-shot search i.e users retrieve information by building a question/query through only a text field and receive the answer in response.…”
HAL is a multi-disciplinary open access archive for the deposit and dissemination of scientific research documents, whether they are published or not. The documents may come from teaching and research institutions in France or abroad, or from public or private research centers. L'archive ouverte pluridisciplinaire HAL, est destinée au dépôt et à la diffusion de documents scientifiques de niveau recherche, publiés ou non, émanant des établissements d'enseignement et de recherche français ou étrangers, des laboratoires publics ou privés.
“…Therefore, there is a need for evaluating the quality of the user search experience when searching for information. And this is especially necessary because of the emergence and popularity of new search techniques such as Elastic search [16] [26], Question Answering over free text [3] [6] [4] [2], and Question Answering over knowledge graphs [8] [9] [27] [30]. The search techniques are one-field one-shot search i.e users retrieve information by building a question/query through only a text field and receive the answer in response.…”
HAL is a multi-disciplinary open access archive for the deposit and dissemination of scientific research documents, whether they are published or not. The documents may come from teaching and research institutions in France or abroad, or from public or private research centers. L'archive ouverte pluridisciplinaire HAL, est destinée au dépôt et à la diffusion de documents scientifiques de niveau recherche, publiés ou non, émanant des établissements d'enseignement et de recherche français ou étrangers, des laboratoires publics ou privés.
“…The records cover a wide range of clinical knowledge, from individual-level information to group-level insight, in various forms, including tables, text, and images [16,17,11,15]. As a vast and comprehensive knowledge base, hospital staff, including physicians, nurses, and administrators, constantly interact with EHRs to store and retrieve patient information to make better clinical decisions [32,3].…”
We present a new text-to-SQL dataset for electronic health records (EHRs). The utterances were collected from 222 hospital staff, including physicians, nurses, insurance review and health records teams, and more. To construct the QA dataset on structured EHR data, we conducted a poll at a university hospital and templatized the responses to create seed questions. Then, we manually linked them to two open-source EHR databases-MIMIC-III and eICU-and included them with various time expressions and held-out unanswerable questions in the dataset, which were all collected from the poll. Our dataset poses a unique set of challenges: the model needs to 1) generate SQL queries that reflect a wide range of needs in the hospital, including simple retrieval and complex operations such as calculating survival rate, 2) understand various time expressions to answer time-sensitive questions in healthcare, and 3) distinguish whether a given question is answerable or unanswerable based on the prediction confidence. We believe our dataset, EHRSQL, could serve as a practical benchmark to develop and assess QA models on structured EHR data and take one step further towards bridging the gap between text-to-SQL research and its real-life deployment in healthcare.36th Conference on Neural Information Processing Systems (NeurIPS 2022) Track on Datasets and Benchmarks.
scite is a Brooklyn-based organization that helps researchers better discover and understand research articles through Smart Citations–citations that display the context of the citation and describe whether the article provides supporting or contrasting evidence. scite is used by students and researchers from around the world and is funded in part by the National Science Foundation and the National Institute on Drug Abuse of the National Institutes of Health.