Autre~57 · IA en attente
I had Claude Code hand off big summarizing jobs to the on-device model in macOS 27
r/ClaudeAIu/TigerKR1 octobre 2026
Analyse IA en cours de préparation : les informations ci-dessous proviennent de la détection automatique.
Résumé
macOS 27 ships with fm, a command-line tool for Apple's on-device Foundation Model. I've been using it to cut Claude Code token use on big piles of text like session transcripts, logs and long documents. The local model writes a summary and Claude reads that instead of the raw text. It all runs on the Mac, so nothing…
Afficher le post originalMasquer le post original
macOS 27 ships with fm, a command-line tool for Apple's on-device Foundation Model. I've been using it to cut Claude Code token use on big piles of text like session transcripts, logs and long documents. The local model writes a summary and Claude reads that instead of the raw text. It all runs on the Mac, so nothing gets sent anywhere.
Here's the prompt I used to have Claude Code build the skill. I spent some time testing fm first, and the prompt includes what I measured, so the skill knows what the small model is good at and where it falls over.
Setup first. You need macOS 27, and you have to run sudo fm license once in Terminal to accept Apple's terms (until you do, every other fm command just exits). After that, fm available should say "System model available".
Paste this into Claude Code:
```text Create a user-level Claude Code skill named "on-device-summarize" at ~/.claude/skills/on-device-summarize/SKILL.md that tells Claude when and how to use Apple's on-device language model (the "fm" command-line tool) to summarize text instead of reading it with Claude tokens. User wants this to save Claude tokens where doing so costs no accuracy. Before writing, confirm the current skill file format (frontmatter fields, how the description triggers the skill) in the Claude Code documentation; use the claude-code-guide agent. Put "effort:" in the frontmatter only if needed (on most models, an effort change reloads the prompt cache). Never put "model:" in it: a skill that names a different model is a model switch, and the next request re-reads the whole conversation with no cache hits.
Intent - Use the on-device model to compress large volumes of text (session transcripts, logs, long documents, exports) into short summaries that Claude then reads, when the raw text would cost many Claude tokens and a faithful summary is enough for the task. - Do not use it where the answer needs judgment across many documents, exact values, or a decision. Claude does that part, reading the summaries plus the authoritative sources. - Accuracy outranks token savings (User's priority order: accuracy, reliability, CPU, memory, time, tokens). When unsure whether a summary is faithful enough, check a sample against the source before relying on the rest.
Tool facts (measured on the Mac, 2026-10-01; re-measure if macOS changes) - Binary: /usr/bin/fm ("Apple Foundation Models CLI"). "fm available" prints "System model available" when it is usable. Commands: available, chat, count-tokens, license, respond, schema, serve. Read "fm <command> --help" for options; strip ANSI color codes from its output. - Summarize: fm respond --no-stream --greedy -i '<instructions>' < file. --greedy makes the output repeatable for the same input. --no-stream returns the whole answer at once. - Count tokens before sending: fm count-tokens -q -i '<instructions>' < file prints a bare integer. - Structured output: fm respond --schema <file>, with the schema built by "fm schema object". For an array of objects, "--array" must come AFTER the "--schema" that defines the object (--object topics --schema "$(fm schema object --name Topic --string name)" --array); placed after --object it is refused.
Limitations (measured) - Context window: an input of 7,103 tokens worked; 8,624 failed with "Error: The session's transcript exceeded the model's context size." The output counts against the same window. Keep each input at or below about 6,500 tokens and chunk anything larger, splitting on natural boundaries (messages, sections, paragraphs), never mid-sentence. - Tokenizer: about 4.3 characters per token for English technical text. - Speed: a 2,500-token input summarized in 11.7 seconds; two calls run in parallel finished in 15 seconds together, so parallelism helps only modestly. Budget roughly 10 to 25 seconds per call and measure the real rate on the first few calls of a large run. - Summary quality: a one-paragraph summary of a 10,700-character handoff was accurate, and its one specific claim checked out word for word against the source. Faithful compression is its strength. - Judgment quality is poor. Asked to list sub-projects with a status, it split one piece of work into six, called unfinished work "finished", and returned a status value outside the list its schema description allowed. A schema guarantees the JSON shape, not the content: descriptions are hints, not enforced values (the schema builder has no enum). - Free-form JSON (no --schema) has no fixed shape between calls, and "extract unique entities" returns generic named things (people, concepts), not task-level items. Do not parse it in a pipeline. - It is a small model: it can drop minor topics when one call covers several unrelated texts. Summarize one coherent unit per call (one session, one document, one chunk) rather than packing unrelated texts together. - Runs only on this Mac, locally; nothing leaves the machine, which makes it suitable for personal content User keeps local.
What the skill must instruct 1. Decide first: is the bulk text only needed in compressed form, and is it large enough (roughly over 20,000 Claude tokens) to be worth the local run time? If not, read it directly. 2. Check "fm available"; if the model is unavailable, say so and fall back to reading with Claude. Never fail silently. 3. Count tokens, chunk to the limit, summarize each unit with --greedy, and keep a map from each summary back to its source (file, line range or message identifier) so Claude can open the original when a summary matters. 4. Treat summaries as leads, not evidence: any claim that goes into a decision, a document or a status is verified against the source. 5. For any run over a handful of calls, validate first: summarize a small sample both ways (fm and Claude reading directly), compare, and report the comparison before running the rest. 6. Report every failed call (the file and the error text) and the totals (calls, failures, elapsed time); a run with failures is not reported as complete. 7. Use Python (python3) for loops and chunking. Put scripts in a temporary scratch directory, not in the project.
Deliver the SKILL.md, lint it with markdownlint if it is installed, and show the final description line so the user can judge when it will trigger.
Remind the user that they will need to have installed macOS 27, and that they will need to run "sudo fm license" in the terminal once, before they can use the fm command. ```
What you end up with:
The skill only kicks in when there's a lot of text (roughly 20k+ Claude tokens) and a summary is good enough for the job. Anything smaller, Claude just reads.
It won't use fm for anything that needs exact values, judgment across several documents, or a decision. The model is good at faithful summaries and bad at judgment. When I asked it for a status list of sub-projects, it split one piece of work into six and marked unfinished work as finished.
It checks that the model is available (and falls back to Claude if not), splits input into chunks of about 6,500 tokens because the context window ran out somewhere between 7,100 and 8,600, uses --greedy so the output is repeatable, and keeps track of where each summary came from so Claude can go back to the original.
Summaries are treated as leads, not facts. Before a big run it summarizes a small sample both ways, with fm and with Claude, and shows you the comparison. Failed calls get reported instead of quietly skipped.
On my Mac a 2,500-token chunk took about 12 seconds. Two calls in parallel finished in 15 seconds, so parallel runs help a bit but not much.
If you try it, I'd like to hear how it does on your machine.