Band, SaurabhSaurabhBand2026-07-232026-07-232026-07-09https://media.suub.uni-bremen.de/handle/elib/25098https://doi.org/10.26092/elib/6245The increasing adoption of Internet of Things (IoT) systems based on resource-constrained sensor nodes for large-scale monitoring has introduced significant challenges in ensuring the reliability of Wireless Sensor Network (WSN). Faults caused after deployment lead to degraded performance or undetected anomalous behavior. Detecting such subtle deviations and diagnosing their root causes remains challenging due to limited on-node resources and lack of transparency into internal system behavior. This thesis addresses these challenges by proposing an integrated framework for improving WSN reliability at the edge layer. At its core, the framework employs an event-trace-based monitoring approach that captures the internal execution behavior of sensor nodes. Based on this, a lightweight runtime fault detection method, VarLogger, is introduced. VarLogger monitors execution traces on-node and detects faults using spatial and temporal features, achieving strong detection performance (F1-score above 0.9) while maintaining low computational and energy overhead. To complement fault detection, this thesis also proposes VarDiag, a fine-grained fault diagnosis method that localizes faults to specific execution subsequences and corresponding source-code locations. VarDiag significantly reduces debugging effort by providing actionable insights, achieving up to 0.97 F1-score and substantially outperforming the state-of-the-art approach. Both methods are evaluated on real hardware across three IoT applications and multiple fault types, demonstrating effectiveness in detecting diverse faults without disrupting system functionality. To address scalability and adaptability, an automated code instrumentation tool is developed and evaluated using rule-based and Large Language Models (LLM)-based approaches, where the LLM-based method shows higher adaptability to new software platforms. In addition, system-level reliability is addressed through redundancy modeling. A Continuous-Time Markov Chain (CTMC)-based model is developed to analyze the trade-off between redundancy and system availability under realistic assumptions. The results show that introducing a single backup device significantly improves availability. Overall, the results demonstrate that lightweight, execution-based monitoring combined with data-driven analysis and system-level design provides an effective and practical solution for improving reliability in IoT systems.enhttps://creativecommons.org/licenses/by/4.0/Fault DetectionNovelty DetectionAnomaly DetectionWSNIoTTime Series AnalysisEvent TracesReliability600 Technology::620 EngineeringA Framework for Runtime Fault Detection and Diagnosis for Reliable Provision of Resource-Constrained IoT DevicesDissertation10.26092/elib/6245urn:nbn:de:gbv:46-elib250982