Title A large-scale llm-driven system for extracting claims about longevity factors and assessing evidence quality from scientific literature
Translation of Title Dideliais kalbos modeliais paremta sistema teiginiams apie ilgaamžiškumo veiksnius ir jų įrodymų kokybei vertinti mokslinėje literatūroje.
Authors Stučinskas, Arnas
Full Text Download
Pages 76
Keywords [eng] longevity, large language models, evidence quality, natural language processing, biomedical literature, claim extraction
Abstract [eng] This thesis addresses the difficulty of comparing scientific claims about longevity factors across a rapidly growing and methodologically heterogeneous biomedical literature. The aim was to develop and evaluate a reproducible computational system for retrieving longevity-related publications, extracting individual factor-outcome claims, checking claim support against source abstracts, and assigning evidence quality grades. A broad PubMed retrieval identified 108,431 records, of which 28,721 were retained after relevance screening. Structured extraction, claim eligibility filtering, entity normalisation, polarity correction, claim validation, hallmark mapping, and evidence tiering produced 2,987 quality-graded claim units from 1,797 publications. The final corpus covered 793 normalised factor nodes and 769 normalised endpoint nodes. The results showed that human longevity research is broad but shallow: only 7.6% of claims reached the highest evidence tier, only 1.1% directly targeted survival or longevity outcomes, and 91.8% of merged factor-outcome claim pairs were supported by a single study. The thesis concludes that claim-level decomposition combined with evidence tiering can make longevity evidence more transparent, comparable, and reusable, while also exposing major gaps in direct and replicated human evidence.
Dissertation Institution Vilniaus universitetas.
Type Master thesis
Language English
Publication date 2026