Item

Toward Usable Scientific Analogies: A Modular Multi-Agent Pipeline for Generation, Explanation, and Evaluation

Barakat, Mariam Amr Hassanien
Citations
Altmetric:
Department
Natural Language Processing
Embargo End Date
Type
Thesis
Date
2026
License
Language
English
Collections
Research Projects
Organizational Units
Journal Issue
Abstract
Analogies support learning by linking new concepts to prior knowledge, yet generating high-quality ones remains challenging due to the need for appropriate sources, structural alignment, and clear explanations. This work explores the problem of automatically generating educational analogies -- specifically, how large language models can identify suitable analogies for unfamiliar concepts, construct meaningful structural mappings, and produce explanations that support learning without human intervention. This research lies at the intersection of educational natural language processing and analogical reasoning, where the goal is not only to generate text but to preserve relational structure in a way that facilitates understanding across domains. Despite recent progress, current research in analogy generation remains fragmented and limited in scope. Existing approaches typically address individual components -- such as retrieval, mapping, or explanation -- in isolation, rather than as part of a unified process. Furthermore, most studies focus primarily on GPT-based models, without systematically exploring a diverse set of large language models or investigating how different input representations affect performance. Evaluation also remains a major challenge, as existing methods rely either on surface-level metrics or on LLM-based judgments that are not consistently validated against human evaluation. To address these limitations, this thesis proposes a modular pipeline for educational analogy generation, grounded in Structure Mapping Theory. The pipeline decomposes the task into four stages: source finding, sub-concept generation, explanation generation, and evaluation. This design enables fine-grained analysis of each stage and allows for systematic exploration of different input configurations. We conduct a comprehensive cross-model evaluation across 12 large language models and 7 embedding models, and evaluate performance on two datasets, SCAR and ParallelPARC. In addition, we introduce a multi-level evaluation framework that combines closed-to-open evaluation, LLM-as-a-judge with multiple judge models, human evaluation, and semantic and distributional analyses. Our results show that optimal models and configurations vary across pipeline stages. Grok-4-Fast performs best for sub-concept matching without background, while Llama-3.1-70B with background leads sub-concept generation, and GPT-4.1-mini achieves the strongest explanation quality. For retrieval, OpenAI's embedding-large with sub-concepts performs best in the closed setting, whereas in the open setting, a Llama-3.1-405B configuration with sub-concepts proves more effective. Sub-concept grounding consistently improves retrieval and explanation quality but offers limited benefit in open-ended source generation. In contrast, Claude Sonnet 4.6 aligns better with human rankings than absolute scores, suggesting that an LLM-as-a-judge is more reliable for comparative evaluation. Additionally, valid analogies are shown to occupy a mid-level range of similarity rather than extreme semantic proximity. Overall, this thesis contributes a modular framework, a multi-level evaluation methodology, and empirical insights into intermediate representations and model diversity, advancing more interpretable and reliable analogy generation for education.
Citation
Barakat, Mariam Amr Hassanien, "Toward Usable Scientific Analogies: A Modular Multi-Agent Pipeline for Generation, Explanation, and Evaluation," M.S. Thesis, Natural Language Processing, MBZUAI, 2026.
Source
Conference
Keywords
Analogical Reasoning, Analogy Generation, Educational NLP, Large Language Models, LLM-as-a-Judge
Subjects
Source
Publisher
DOI
Full-text link