PofoliaShared via Pofolia

ACM Transactions on Software Engineering and Methodology· 2026Q1

“Wait, My Tool Can’t Do That?” A Fine-Grained Capability Study of Modern Static Taint Analyzers

Yanjie Zhao, Shenao Wang, Jian Zhao, Junjie He et al.

Short summary

A new benchmark, xAST, reveals that even leading static taint analyzers only correctly identify expected taint flows (positive cases) and suppress false alarms (negative cases) 61.7% of the time, with missed flows being the primary failure mode.

AI-generated from the title and abstract; the full text is not read.

Key points

  • The xAST benchmark contains 844 paired capability cases across Python, Go, Java, and JavaScript.
  • Eight mainstream static taint analyzers were evaluated on the xAST benchmark.
  • The highest pass rate for correctly detecting flows and suppressing false alarms was 61.7% (CodeQL on JavaScript).
  • Missed expected flows (false negatives) accounted for 36.1% of all failures.
  • Recurring weaknesses include field/element precision, path feasibility, concurrency, and advanced language idioms.

AI-generated from the title and abstract; the full text is not read.

Abstract

Static taint analyzers are widely used in software security, yet aggregate vulnerability-detection metrics provide limited insight into the specific language constructs and analysis capabilities that cause tools to succeed or fail. This paper presents xAST , a fine-grained benchmark containing 844 paired capability cases, implemented as 1,825 individual source files across Python, Go, Java, and JavaScript. The benchmark covers 16 capability dimensions and 76 sub-categories, including context, flow, path, field, and element precision, as well as functions, modules, concurrency, expressions, and dynamic features. Each pair contains a positive instance in which an expected source-to-sink flow should be detected and a structurally related negative instance in which the corresponding alarm should be suppressed. We evaluate eight mainstream static taint analyzers under 14 tool-language configurations. The highest pair pass rate is 61.7%, achieved by CodeQL on JavaScript. Across all 2,881 tool-language pair evaluations, 47.5% pass both instances, 36.1% are false-negative-only failures, 14.3% are false-positive-only failures, and 2.1% fail both instances. The aggregate T-instance detection rate is 61.9%, while the aggregate F-instance suppression rate is 83.6%, showing that missed expected flows are the dominant global failure mode, although individual tools exhibit substantially different detection and suppression profiles. Fine-grained analysis further identifies recurring weaknesses in field and element precision, path-feasibility reasoning, transfer-rule coverage, concurrency and asynchronous modeling, advanced language idioms, alias analysis, and module resolution. These results provide a capability-oriented diagnostic complement to application-level and vulnerability-level benchmarks and offer actionable guidance for tool developers, researchers, and practitioners.

The authors' abstract, as published at the source. ACM Transactions on Software Engineering and Methodology, 2026 · DOI ↗

TakeawaysPremium
Ask the paperFree account

Continue with a free account

Ask the paper: 3 free questions a day about this paper; save it, get its citation, new summaries every day for your field. Takeaways are Premium.

Continue free on the web

Sign in with Google or Apple; no card needed. You come back to this paper.

On your phone:

Field: Software

SoftwareComputer Science