METACURATE-ML: Generalising the extraction of questionnaires to DDI-Lifecycle

Chandresh Pravin; Suparna De; Jon Johnson

doi:10.5281/zenodo.18182866

Back

METACURATE-ML: Generalising the extraction of questionnaires to DDI-Lifecycle

Conference presentation

Open access

METACURATE-ML: Generalising the extraction of questionnaires to DDI-Lifecycle

Chandresh Pravin, Suparna De and Jon Johnson

Annual European DDI User Conference (EDDI2025), 17th (Budapest, Hungary, 02/12/2025)

02/12/2025

DOI: https://doi.org/10.5281/zenodo.18182866

Abstract

natural language processing

Information Extraction

Multimodal layout analysis

Page object detection

Machine Learning

The canonical source for most large scale surveys remains for the past and the foreseeable future the ‘paper’ version of the questionnaire as a PDF. These contain the essential information on the fielding of the survey, the questions asked including response options, ordering and the logic associated with filtering to create DDI-Lifecycle metadata. However, there is no standard layout and formatting in these source documents which limits the scalability of heuristic programming approaches on OCR’ed text.

The presentation will outline an approach to creating DDI-Lifecycle from semi-structured text in social science questionnaires using a combination of text-layout large language models (LLMs) and knowledge graph construction approaches, including methods from digital signal processing.

Files and links (1)

pdf

EDDI-2025-CP-MetaCurate-ML final1.04 MBDownload View

SlidesCC BY V4.0, Open Access

Metrics

1 Record Views

Details

Title: METACURATE-ML: Generalising the extraction of questionnaires to DDI-Lifecycle
Creators: Chandresh Pravin (Corresponding Author)
Suparna De (Author) - University of Surrey, School of Computer Science & Electronic Engineering
Jon Johnson (Author) - University College London
Conference: Annual European DDI User Conference (EDDI2025), 17th (Budapest, Hungary, 02/12/2025)
Grants: Metadata uplift in Longitudinal Population Studies (provenance, discovery and disclosure), UKRI2700, Engineering and Physical Sciences Research Council (United Kingdom, Swindon) - EPSRC
Identifiers: 991082928802346
Academic Unit: School of Computer Science & Electronic Engineering
Language: English
Resource Type: Conference presentation

METACURATE-ML: Generalising the extraction of questionnaires to DDI-Lifecycle

Abstract

Files and links (1)

Metrics

Details

Usage Policy