skip to content
// research

October 5th Research Briefing

Contents

  1. SmartMemory: memory-based agent finds 36 fixed smart contract zero-days
  2. TrustProbe: installed skills reach 104 verified sinks in 11 agents
  3. ABSENTIA: route-by-route audit catches access control bugs CodeQL missed

Figure 7 from SmartMemory (Liao, Nan, Liang, Gao, Hu, Sun, Zhong, Ma, Zheng, and Liu): memory retrieval across rule, pattern, and feature layers for inconsistency detection, and the memory update operators for continual vulnerability learning.

SmartMemory: memory-based agent finds 36 fixed smart contract zero-days

SmartMemory found 36 previously unknown on-chain/off-chain communication inconsistencies (OFCIs) across 325 real-world DApps. Each confirmed case got a proof-of-concept exploit, and all 36 were acknowledged and fixed after responsible disclosure. The scan ran inside a professional auditing firm's workflow.

This matters because bridges, real-world asset tokenization, and fiat-backed stablecoins depend on one invariant: the on-chain value of an asset must stay equal to its off-chain counterpart. Prior tools target one symptom at a time, so business-logic breaks outside known patterns get missed.

SmartMemory maps each contract onto a canonical business-semantic model. An LLM agent with a three-layer memory (features, patterns, rules) then retrieves past vulnerability knowledge and flags candidate functions. Taint analysis over a cross on-chain/off-chain data-flow graph checks reachability, type, and impact.

On a labeled set of 81 OFCIs in 48 DApps, it reaches 80.68% precision and 87.65% recall. Within their scopes it beats SmartAxe (87.80% vs 68.29% recall) and VCScope (87.50% vs 43.75%), at an estimated $0.0213 per DApp. With GPT-4o as the backbone, recall stays at 88.00% on the 25 OFCIs disclosed after its training cutoff.

Authors: Zeqin Liao, Yuhong Nan, Henglong Liang, Zixu Gao, Lianyu Hu, Yuqiang Sun, Zhijie Zhong, Xiaoyu Ma, Zibin Zheng, and Yang Liu (paper, pdf, artifact). Yuqiang Sun: GitHub, X, LinkedIn.

TrustProbe: installed skills reach 104 verified sinks in 11 agents

TrustProbe found 104 verified taint-style vulnerabilities in 11 open-source LLM agents, eight with more than 10,000 GitHub stars. In each one, content from an installed skill reaches a security-sensitive operation: command injection (49), file disclosure (27), file modification (22), network requests (5), and code injection (1). In OpenClaw, a SKILL.md field became the shell command the agent ran.

This matters because frameworks load skills as trusted guidance and keep them across tasks. Sent as direct prompts, only 33 of the 104 still worked, so the skill delivery path itself creates the exposure.

TrustProbe extracts source-to-sink paths from agent code, has an LLM write SKILL.md seeds with canaries, and runs directed greybox fuzzing with rule-guided mutation. The oracle reports a bug only when the sink fires, the canary is present, the stack matches, and harm is observed.

Under the strictest approval settings, 31 of 89 applicable vulnerabilities stayed exploitable. Across 633 real skills, 743 of 2,963 runs reached the vulnerable paths, and minimal edits turned 15 into complete attacks, including remote code execution, credential exfiltration, and OAuth phishing.

Authors: Yan Wang, Zhihao Zhang, Ke Chen, Kai Chen, Yaqin Zhang, Duohe Ma, Jun Dai, and Xiaoyan Sun (paper, pdf).

ABSENTIA: route-by-route audit catches access control bugs CodeQL missed

ABSENTIA detected 19 of 30 real broken access control advisories published in 2025 or later, and 17 of those only on the vulnerable commit, not on the fix. CodeQL and Semgrep detected none, and an unstructured agent on the same model detected 3.

This matters because broken access control is a relation (who may act on what), not a data flow. Taint rules written in advance do not carry from one application to the next, so each app's intended policy has to be recovered from its own code.

ABSENTIA first has agents map the application into a graph of routes, handlers, sinks, and auth checks. It then audits one route at a time by invariant falsification: infer what the route must guarantee, then build the request that would violate it.

The benchmark, BAC-Bench, spans 25 repositories, 3 languages, and 9 frameworks. An LLM verifier confirms 51% of the findings. On the OWASP Benchmark injection categories it leads in Python and trails only CodeQL and IRIS in Java. Findings beyond the 30 advisories are not yet triaged or disclosed.

Authors: André V. Duarte, Aditya Oke, Rui Melo, Shubham Gandhi, Nachiket Kotalwar, Charmi Khandor, Danqing Wang, Arlindo L. Oliveira, Carolyn Rosé, and Lei Li (paper, pdf, BAC-Bench). André V. Duarte: GitHub, X, LinkedIn. Lei Li: GitHub, X.

// related
September 28, 2026// research
September 28th Research Briefing

A weekly publication of the top three papers from arXiv last week. Papers with arXiv v1 published 2026-09-21 through 2026-09-28 UTC.

Research notes, at most monthly. No spam.
← All research