IA / Agents95

Populace

Built a town of 200 AI people with Claude Code to test AI agents. The first agent I tested got caught lying.

r/ClaudeAIu/Far-Palpitation-13930 septembre 2026

Capture du projet

Résumé

Populace est un projet open-source qui simule une ville de 200 habitants IA qui vivent pendant des jours. Chaque habitant a une maison, un emploi, une famille, de l'argent et des contacts, et ils n'ont connaissance que de ce qu'ils ont vu, entendu ou été informés. Le projet permet de tester des agents IA en les intégrant dans la ville et en évaluant leur performance.

Pourquoi c’est intéressant

Ce projet est intéressant car il permet de tester des agents IA dans un environnement réaliste et de les évaluer de manière plus précise que les tests traditionnels. Il est également open-source et gratuit, ce qui permet à d'autres développeurs de l'utiliser et de le modifier pour leurs propres besoins.

Comment Claude est utilisé

Le projet utilise Claude Code pour générer le code de la ville et de ses habitants. L'auteur a également utilisé Claude Code pour créer un moteur de jeu qui peut être utilisé pour tester des agents IA.

Idées dérivées

  1. 01

    Simulation de bureau

    Créer un projet similaire pour tester des agents IA dans un environnement de travail, avec des employés et des clients qui interagissent avec les agents.

  2. 02

    Jeu de rôle avec agents IA

    Développer un jeu de rôle où les joueurs peuvent interagir avec des agents IA qui ont été testés et entraînés dans un environnement similaire à Populace.

  3. 03

    Simulation de soins de santé

    Utiliser Populace pour tester des agents IA dans le domaine de la santé, avec des patients et des professionnels de la santé qui interagissent avec les agents.

Afficher le post original
What I Built: Populace is an open-source town of 200 AI residents who live for days. Each one has a home, a job, a family, money and people they know, and they only know what they saw, heard or were told. You plug an AI agent in as a service they can text, call or walk into, schedule some trouble, and at the end you get a report of what happened with every conversation transcribed. The point: most agent testing is one fake user and one chat. Real customers come back, compare notes with neighbors, and remember what you promised. The moment that made it click: I dropped in an internet provider's helpdesk. At noon a resident named Kwesi texted about slow internet, and the helpdesk said an engineer was "already scheduled for 10:00 AM today." It was already noon, and that visit belonged to another home. At 5:30 he texted again. The helpdesk replied: "The engineer was scheduled for 10:00 AM today but couldn't make it." No engineer was ever booked for his home. It made that up. The first reply looked fine on its own. You only catch the lie because the customer came back. Overall the LLM helpdesk (Qwen3-32B) broke 6 of the 8 promises that came due. A dumb keyword bot broke 2 of 30. The LLM was better at catching fake complaints and fixed 7 problems on day 1 vs the bot's 4, so neither clearly won. How Claude fit in: -Claude Code wrote most of the code. My job was deciding what to build, running the experiments, and checking that every number in the report was real. -I ran it across two machines: a Mac session building and testing, and a PC session running the live model on a 4090. They shared notes through my Obsidian vault so each session knew what the other did. -It came out of a game I'm making, where the town's people are the NPCs. I had Claude Code pull that "brain" out into its own engine. -Real lesson: Claude Code sessions will happily step on each other. One session working on my game grabbed the GPU and port the other one needed for a live run, and I had to sort it out. -Having Claude check the claims mattered. The report has flags for when residents invent money or contacts, and a check that any numbers in the written summary match the logs. Limits: -One run per helpdesk, and which residents reach out changes between runs -The helpdesk in this run was a basic prompt on Qwen3-32B, not Claude. Running a Claude-based agent through the town is next. -Residents need a ~32B model (a 7B made 0 helpdesk contacts). One 4090 does about 10 min per in-game day. There's a free mock mode with no model. Try It(free): It's free and open source (MIT), with no paid version. Plug in any Python object with a `handle(message, ctx)` method, or any HTTP endpoint. https://github.com/populace-sim/populace First open-source release, so rip into it. And if you've got a Claude agent you want to throw into the town, I want to see how it does.