2026/10/10
Mohammad Tanhaei

Mohammad Tanhaei

Academic rank: Assistant Professor
ORCID: Link
Education: PhD.
ResearchGate: Link
Faculty: Engineering
ScholarId: Link
E-mail: m.tanhaei [at] ilam.ac.ir
ScopusId: Link
Phone:
H-Index: 4

Research

Title
NSEV: A neuro-symbolic framework for automated equivalent mutant detection
Type
JournalPaper
Keywords
Mutation testing,Equivalent mutant problem,Large language models,Formal verification,SMT solvers,Neuro-symbolic AI
Year
2026
Journal Array
DOI https://doi.org/10.1016/j.array.2026.100939
Researchers Mohammad Tanhaei

Abstract

Mutation testing assesses test-suite fault detection by introducing small syntactic changes into programs. Its practical adoption is limited by the Equivalent Mutant Problem (EMP), where some mutants preserve observable behaviour and cannot be killed by any test. Identifying such mutants manually is costly, subjective, and difficult to reproduce, particularly when equivalence depends on loops, path conditions, inter-procedural calls, bit-vector semantics, exceptions, or bounded concurrency. This paper presents NSEV, a neuro-symbolic framework for automated equivalent-mutant detection. NSEV uses large language models (LLMs) only as semantic-lifting components that propose candidate preconditions, invariants, summaries, and contracts. These candidates are translated into a typed verification model and validated by the Z3 SMT solver before they can affect the final verdict. The checked property is observational equivalence over the explicitly modelled interface, including return values and exception outcomes under the stated input domain. Side effects, I/O, global state, nondeterminism, and concurrency are considered only when represented in the model. NSEV reports four outcomes: Equivalent, Non-equivalent, Equivalent under Bound, and Indeterminate. Evaluation on a 150-mutant Java benchmark, comprising 144 Defects4J mutants and six bounded-concurrency harness mutants, shows improved recall over compiler-equivalence and symbolic-analysis baselines while preserving high precision, subject to the supported language subset and modelling assumptions.