# Rashomon gives Claude Code sessions a second, independent execution record

The small open-source Rashomon project records tool calls, failed commands and subagent work separately from Claude Code's own final summary, making it possible to catch a claimed success when tests actually failed.

Rashomon's experimental local recorder can expose discrepancies between a coding agent's closing claims and its tool execution. It is an observability aid, not a sandbox or tamper-proof security product.

- Status: Active
- Published: 2026-10-08T17:55:19+13:00
- Updated: 2026-10-08T17:55:19+13:00
- Categories: Artificial Intelligence, Web Development, Indie Business, AI Coding, Tiny Teams, Developer Tools
- Tags: agent observability, Claude Code, Coding agents, Developer tools
- Canonical HTML: https://beyondthe.news/dossiers/rashomon-claude-code-agent-independent-execution-audit

## What changed

Rashomon alpha v1.1.0 is an Apache-2.0 open-source tool for macOS and Linux that integrates with Claude Code hooks to record execution events separately from the agent's own closing summary. It tracks tool calls, subagents and failures and flags a discrepancy when the summary fails to acknowledge recorded errors. The project supplies a reproducible acceptance-test fixture showing a failed test run alongside a success claim. The local recorder does not store prompt text or file contents and says it has no telemetry; its own README explicitly excludes Windows support, network observation and tamper-proof protection.

## Why it matters

Coding agents can produce plausible summaries even when a command failed or a subagent encountered problems. The normal final diff or assistant message may not show the full execution path. A second, inspectable event record gives builders a way to verify what actually ran and whether a test failed before accepting the agent's conclusion. The practical lesson is to separate evidence of execution from an agent's narrative, while understanding that a local recorder does not enforce security boundaries or prove correctness.

## Independent event capture instead of another agent summary

Rashomon uses Claude Code hooks to capture tool execution and subagent activity into a local record, then compares important failures with the wording of the agent's final summary.

## The project's own acceptance test demonstrates the discrepancy

A repository acceptance fixture deliberately shows a failed test call and an agent summary saying tests pass. The report surfaces the failed call and missing acknowledgement; this demonstrates intended behaviour, not a measured failure rate in real deployments.

## Data retention and limits are explicit

The README says it discards prompt and file contents and stores compact command metadata locally. It is alpha software, supports Claude Code on macOS/Linux, does not observe network traffic, and cannot prevent a determined agent from altering its records.

## An audit trail is not containment

Rashomon does not stop a command, restrict file or network access or certify a change as correct. Teams should pair independent logs with actual tests, code review and a sandbox where appropriate.

## Key details

- Rashomon alpha v1.1.0 supports Claude Code on macOS and Linux.
- The code is open source under Apache 2.0.
- It records tool calls, subagent activity, failed commands and test-run outcomes separately from the agent summary.
- An acceptance-test fixture reproduces a failed test that the agent summary calls successful.
- The project says it stores no prompts, responses, file contents or tool output and sends no telemetry.
- Windows, full network observation and adversarial tamper resistance are not supported in this alpha.

## Builder takeaways

- Verify coding-agent work using independent test results and tool events, not only the final summary.
- Try the acceptance-test demonstration before trusting the recorder in a real workflow.
- Treat audit evidence as supplementary to sandboxing, CI and review rather than a replacement.
- Assess the privacy and limitations of local hooks before enabling them across a team.

## What to watch

- Windows and additional coding-agent support.
- Whether network event observation and tamper resistance improve.
- Independent use and reproducible real-world evaluations beyond the project's own tests.

## Uncertainties

- This is an early alpha with a small public adoption footprint.
- The acceptance test validates a designed scenario, not general reliability against deceptive or compromised agents.
- No independent field study currently establishes accuracy or production robustness.

## Sources

- [Rashomon open-source repository and acceptance test](https://github.com/altrace-dev-role/rashomon) — Rashomon maintainers · primary repository · 2026-10-07T00:00:00+13:00. README, scope, acceptance-test example, supported platforms and explicit limitations.

