IndusTSLM: Exploring Time-Series Language Models for Drilling Data
Shaaban, Yahia Salaheldin
Shaaban, Yahia Salaheldin
Citations
Altmetric:
Author
Supervisor
Department
Machine Learning
Embargo End Date
Type
Thesis
Date
2026
License
Language
English
Collections
Research Projects
Organizational Units
Journal Issue
Abstract
Modern oil and gas drilling has become an increasingly sophisticated process. A single well- bore can generate an average of 700,000 sensor measurements per day. This growing volume of data presents both an opportunity and a challenge. While these time series carry rich operational signatures, manual interpretation is expensive, time consuming, suboptimal when prompt decisions are needed, and prone to overlooking critical events. Meanwhile, despite the rapid rise of large language models and agentic systems, longitudinal sensor data remains one of the major open challenges, as it demands contextual reasoning that traditional foundation models lack. Recent advances in Time Series Language Models (TSLMs) have shown that aligning temporal patterns with the representational power of language models can unlock new capabilities, from automated report generation to natural language grounded complex reasoning. However, most existing work has been developed and evaluated on curated public benchmarks, leaving a significant gap in understanding how these approaches perform on real world industrial data characterized by noise, irregular sampling, long operational horizons, and domain specific semantics. This thesis investigates the alignment between two modalities, drilling sensor time se- ries and textual representations from pretrained language models, through a progressive architectural exploration. We begin with a dual encoder baseline that learns shared em- beddings for time series segments and their textual descriptions, and advance through to the adoption of a Flamingo style architecture that allows the language model to attend to time series features without modifying its pretrained weights. Drawing on the insights from these experiments, we propose IndusTSLM, a model tailored for industrial time series language understanding, leveraging techniques in data curation and training recipe design. We evaluate the framework on proprietary drilling datasets and demonstrate its capacity to bridge the gap between raw sensor streams and human interpretable operational insights. We further introduce DrillBench, a standardized benchmark comprising seven task types across four groups built on publicly available drilling data, with a knowledge decoupling framework that distinguishes genuine sensor analysis from pretrained domain knowledge exploitation. Our findings suggest that grounding language models in real industrial time series is not only feasible but offers a promising path toward intelligent, context aware decision support in drilling operations.
Citation
Shaaban, Yahia Salaheldin, "IndusTSLM: Exploring Time-Series Language Models for Drilling Data," M.S. Thesis, Machine Learning, MBZUAI, 2026.
Source
Conference
Keywords
multimodal alignment, Time series language models, foundation models
