DevOps95

Velme

4 compiler phases in 5 days with Claude Code: the main session never writes code, and a separate reviewer caught every bug that passed tests. ~460M tokens.

r/ClaudeAIu/reypham29 septembre 2026

Capture du projet

Résumé

Le projet Velme est un langage de programmation qui utilise Claude Code pour générer du code, mais avec une approche innovante : un compilateur déterministe et un sandbox décident si le code est accepté, plutôt que de laisser le modèle LLM juger son propre travail. Le projet a été développé en 5 jours, avec 4 phases et environ 21 600 lignes de code Rust.

Pourquoi c’est intéressant

Ce projet est intéressant car il montre comment Claude Code peut être utilisé pour générer du code de haute qualité, tout en maintenant un contrôle déterministe sur le processus de compilation et de sandbox. L'utilisation de sous-agents pour la construction et la révision du code est également une approche innovante.

Comment Claude est utilisé

Développé en grande partie avec Claude Code ; l'application utilise des sous-agents pour la construction et la révision du code, tandis que la session principale coordonne et vérifie les résultats.

Idées dérivées

  1. 01

    Révision de code automatisée

    Créer un système de révision de code automatisé pour d'autres langages de programmation, en utilisant Claude Code pour analyser et améliorer la qualité du code.

  2. 02

    Outil d'aide à la décision

    Développer un outil d'aide à la décision pour les développeurs, en utilisant Claude Code pour analyser les besoins et les contraintes d'un projet et proposer des solutions optimales.

  3. 03

    Formation de développeurs

    Concevoir un système de formation pour les développeurs, en utilisant Claude Code pour créer des exercices et des défis personnalisés pour améliorer leurs compétences.

Afficher le post original
I'm building a programming language in Rust with Claude Code. In 5 days: 4 phases, ~21.6k lines of Rust, 236 acceptance criteria tracked to tests. My first attempt ran everything in one long session. It hit context limits and the code got worse after each /compact. What fixed it: the main session only coordinates. It spawns subagents and reads their short reports, never diffs. The loop for each slice (a slice is about one commit): • Build. A slice-builder subagent gets the slice ID, plan text and the exact spec sections. The model and effort for each slice come from a progress table. • Review. A separate slice-reviewer (Opus, read-only) checks the uncommitted diff. • Fix. Findings go back to the same builder via SendMessage. It still has the context, so fixes are cheap. • Commit. The main session re-runs verification itself. It never trusts "all tests pass" from the builder. Before closing a phase, three reviews run one after another, each fully fixed before the next: architect review (Opus, high) → code review of the whole phase → /simplify. Running them in parallel caused conflicting edits. What the reviews caught after everything already passed tests: • A sandbox cost bug: comparing strings cost the same regardless of length, so generated code could burn huge CPU cheaply. • A cache key that left out some limits, so two different contracts could share cached code. • The artifact store followed symlinks and read unbounded files. • The writer accepted files the reader would refuse, so you could save something you couldn't load. None were caught by the agent that wrote the code. What it cost, in tokens (from my Claude Code transcripts): • ~460M processed; 97% cache reads. • Main session: 39%, even though it writes no code. Every turn re-reads the conversation. • Subagents: 61%. • ~115M per phase. • Opus ~90%, Sonnet ~10%, Haiku 0%. I planned Sonnet builders and Haiku for mechanical work, but kept bumping "hard" slices to Opus. Next phase I'm pushing more to Sonnet to see if review catches the difference. Lessons: • A separate reviewer beats self-review. • Brief agents with rule IDs (INV-4, AC-RDM-08) so they grep instead of reading whole files. • Add one line to every agent prompt: "If it's ambiguous or needs a judgment call, stop and report instead of guessing." • Measure your model mix; the plan and reality drift. What I'm building: Velme (https://github.com/velme-lang/velme), a language where an LLM writes the implementation but a deterministic compiler and sandbox decide whether it's accepted. It's the same idea as the workflow: the LLM never judges its own work. Happy to share the agent files and CLAUDE.md setup in the comments. How do others split work between models?