Retrieval-augmented generation (RAG) systems answer a user query by first retrieving evidence from an external document collection and then asking a language model to generate a response from that evidence. This workflow is useful for domain-specific and citation-oriented applications, but it also creates hidden failure points. A response may look fluent and relevant even when the retrieved evidence is outdated, the needed passage has been truncated, the citation no longer points to the supporting source, or the system has silently used a fallback path without retrieval. We study this deployment-relevant failure mode under the term silent failure: an execution in which a response passes a coarse answer-level acceptability check but violates at least one deeper evidence, citation, abstention, provenance, or operational criterion. We report a controlled offline fault-injection study across three RAG system baselines: a sparse BM25 baseline, a dense BGE-based baseline, and a hybrid BM25+dense system with reranking. The injected faults cover stale indexes, chunk-boundary corruption, embedding mismatch, top-k truncation, reranker disabling, distractor injection, context overflow, context shuffling, prompt drift, citation suppression, fallback bypass, and duplicate-source inflation. The study instruments retrieval diagnostics, context-level measurements, answer-level factuality assessment, citation checks, runtime signals, and candidate gate decisions over 146,016 query execution traces. In this experimental setting, several fault classes preserved answer-level acceptability while degrading evidence support, citation correctness, abstention behavior, or provenance retention. Compared with single-layer detector baselines based on retrieval, faithfulness, citation integrity, or operations alone, the evaluated composite multi-layer gate achieved stronger held-out detector metrics under the selected injected faults. The findings motivate pre-release validation and runtime-observability templates for RAG pipelines, while emphasizing that the reported thresholds are validation operating points and require deployment-specific calibration before production use.