2026/10/10
Mohammad Tanhaei

Mohammad Tanhaei

Academic rank: Assistant Professor
ORCID: Link
Education: PhD.
ResearchGate: Link
Faculty: Engineering
ScholarId: Link
E-mail: m.tanhaei [at] ilam.ac.ir
ScopusId: Link
Phone:
H-Index: 4

Research

Title
Silent-Failure Detection in RAG Pipelines: A Controlled Fault-Injection Study of Multi-Layer Quality Gates
Type
JournalPaper
Keywords
Retrieval-Augmented Generation; Large Language Models; Machine Learning Testing; Fault Injection; Deployment-Oriented Validation; Quality Gates; Silent Failures
Year
2026
Journal Results in Engineering
DOI https://doi.org/10.1016/j.rineng.2026.113085
Researchers Mohammad Tanhaei

Abstract

Retrieval-augmented generation (RAG) systems answer a user query by first retrieving evidence from an external document collection and then asking a language model to generate a response from that evidence. This workflow is useful for domain-specific and citation-oriented applications, but it also creates hidden failure points. A response may look fluent and relevant even when the retrieved evidence is outdated, the needed passage has been truncated, the citation no longer points to the supporting source, or the system has silently used a fallback path without retrieval. We study this deployment-relevant failure mode under the term silent failure: an execution in which a response passes a coarse answer-level acceptability check but violates at least one deeper evidence, citation, abstention, provenance, or operational criterion. We report a controlled offline fault-injection study across three RAG system baselines: a sparse BM25 baseline, a dense BGE-based baseline, and a hybrid BM25+dense system with reranking. The injected faults cover stale indexes, chunk-boundary corruption, embedding mismatch, top-k truncation, reranker disabling, distractor injection, context overflow, context shuffling, prompt drift, citation suppression, fallback bypass, and duplicate-source inflation. The study instruments retrieval diagnostics, context-level measurements, answer-level factuality assessment, citation checks, runtime signals, and candidate gate decisions over 146,016 query execution traces. In this experimental setting, several fault classes preserved answer-level acceptability while degrading evidence support, citation correctness, abstention behavior, or provenance retention. Compared with single-layer detector baselines based on retrieval, faithfulness, citation integrity, or operations alone, the evaluated composite multi-layer gate achieved stronger held-out detector metrics under the selected injected faults. The findings motivate pre-release validation and runtime-observability templates for RAG pipelines, while emphasizing that the reported thresholds are validation operating points and require deployment-specific calibration before production use.