Outil IT95

Receipts

I built a Claude Code skill that makes Claude prove its bug fixes: every test it writes has to fail without the fix

r/ClaudeAIu/Sanechka_SS29 septembre 2026

Capture du projet

Résumé

Le projet Receipts est un plugin Claude Code qui vérifie les tests ajoutés ou modifiés par Claude pour s'assurer qu'ils prouvent réellement les corrections de bogues. Il exécute chaque test deux fois : une fois avec le changement et une fois avec les fichiers source modifiés rétablis à leur version d'origine. Les résultats sont classés en trois catégories : PROVEN, THEATER et WEAK.

Pourquoi c’est intéressant

Ce projet est intéressant car il aborde un problème courant dans le développement logiciel : la vérification de l'efficacité des tests pour prouver les corrections de bogues. La solution proposée est originale et utile, car elle permet de renforcer la confiance dans les corrections de bogues apportées par Claude.

Comment Claude est utilisé

Le projet utilise Claude Code pour développer un plugin qui vérifie les tests ajoutés ou modifiés par Claude, en les exécutant deux fois : une fois avec le changement et une fois avec les fichiers source modifiés rétablis à leur version d'origine.

Idées dérivées

  1. 01

    Amélioration des tests existants

    Développer un outil qui analyse les tests existants pour identifier les cas où les tests pourraient être améliorés pour mieux prouver les corrections de bogues.

  2. 02

    Intégration avec d'autres outils de développement

    Créer un plugin pour d'autres outils de développement qui intègre la fonctionnalité de vérification des tests, permettant ainsi une utilisation plus large de cette technologie.

  3. 03

    Prédiction de la pertinence des tests

    Concevoir un système qui utilise l'apprentissage automatique pour prédire la probabilité qu'un test soit pertinent pour une correction de bogue donnée, aidant ainsi les développeurs à se concentrer sur les tests les plus importants.

Afficher le post original
Claude Code fixes a bug, adds a test, CI goes green, and it says "done". But a green test only says the test passes. It doesn't say the test would have failed before the fix. So I built Receipts. It runs every test a change adds or edits twice: once with the change, and once with the changed source files reverted to main. A test for a fix has to fail without it. • PROVEN: fails without the fix, passes with it. That's the one you want. • THEATER: passes both ways. It would have passed before the fix too, so it proves nothing. • WEAK: fails without the fix only because it imports something the fix added. It ships as a Claude Code plugin with a prove-fix skill. After Claude fixes a bug, it runs the check on its own tests before saying it's done. In a real session, Claude wrote a test for is_leap(2020), got THEATER (2020 was never broken; 1900 was), wrote a test for 1900 instead, and got PROVEN, without touching the fix. /plugin marketplace add syntaxixr/receipts /plugin install receipts-check@receipts To see how often this matters, I ran it over 100 pull requests with coding-agent fingerprints (mostly Claude Code) in Claude Agent SDK, OpenAI Agents SDK, the MCP Python SDK, fastmcp and simonw/llm, plus 81 maintainer fix commits in libraries like click and sqlparse: • 82% of agent PRs and 90% of maintainer fixes were proven. Most tests do their job. • In 10% of agent PRs, every test failed on the old code only because the test file imported a name the PR added at the top. On the old code the file can't even load, so nothing ever ran against the old behavior. No maintainer commit did that. Example: claude-agent-sdk-python #1016 (https://github.com/anthropics/claude-agent-sdk-python/pull/1016). Its test file imports TaskUpdatedMessage and TERMINAL_TASK_STATUSES at the top. On the old code the whole file fails to import, so all 10 tests "fail", including an existing test for unknown subtypes. The fix may well be right; the tests just can't show it. There's also an opt-in Stop hook that won't let Claude end a turn while its changed tests prove nothing. It's off by default, because it runs the repo's tests. No LLM in the check itself: it's just your pytest, vitest or jest, run twice. Repo, GIFs and the full study: https://github.com/syntaxixr/receipts Happy to hear where the verdicts are wrong for your code.