Anthropic said in a developer blog post on Sept. 28 that it has added two commands to the claude-api skill available in Claude Code: /claude-api build-eval, which helps developers create evaluation sets for their own AI applications, and /claude-api hillclimb, which adjusts prompts, models, and parameters step by step based on evaluation results while trying to prevent overfitting.
In an internal customer support case published by Anthropic, Claude changed a setup that originally used Opus 4.8 with high reasoning intensity into a cheaper configuration based on Sonnet 5 with low reasoning intensity. On unseen test tickets during the tuning process, accuracy increased from 78.6% to 90.5%, while cost dropped to about one-fifth of the starting level.
Support evaluation moved from Opus 4.8 to Sonnet 5
The support evaluation included 44 tickets in total. Of those, 30 were used for tuning and 14 were held out for testing. The starting point was Opus 4.8 at its default high reasoning intensity. On the tuning tickets, decision accuracy was 74.4%, and token cost was 4.6 cents per ticket.
Claude first reviewed the prompt and removed a mandatory tool-calling flow, draft steps, and conflicting rules. It then switched to Opus 5.5 with low reasoning intensity. Accuracy rose to 87.8%, and cost fell to 1.9 cents per ticket.
Anthropic said part of the savings came from Opus 5.5 pricing. Input and output tokens were 20% cheaper than Opus 4.8, while cache reads were 60% cheaper.
After Opus 5.5 met the target, Claude tested the cheaper Sonnet 5 at low reasoning intensity. Accuracy came in at about 88.9%, and cost was cut in half again to about 1 cent per ticket. After adding routing rules and a cross-check on refund limits, Sonnet 5 reached 98.9% on the tuning tickets.
hillclimb changes one thing at a time
Before hillclimb starts, Anthropic said Claude asks the developer what should be optimized, such as improving performance or reducing cost without changing performance. It then randomly splits the evaluation set into a tuning group and a test group.
In each round, Claude reads only the failed records in the tuning group and proposes one modification. If the tuning score improves but the test score does not, the system treats that as a sign of possible overfitting and rolls the change back. A change is kept only when both groups improve.
The post gave an example involving image text recognition. If the evaluation questions happen to require OCR, the tuning process may add an OCR tool to the application, lifting the score on the test but not necessarily helping in real use. Anthropic said failed content should not be pasted directly into prompts, and answers should be structured so the model cannot directly access them.
What happens when scores stall
If scores stop improving for two to three rounds in a row, Claude checks the remaining failed cases one by one and groups them by cause. The goal is to identify whether the issue comes from ambiguous questions, a faulty grader, or problems in the evaluation environment.
Anthropic said it used the same process to tune the claude-api skill itself, raising pass rates from 66% to about 88%. During that work, the company found that one grader required something that did not match the question, while another grader included instructions that conflicted with the official documentation.
How build-eval creates an evaluation set
Anthropic said build-eval starts by interviewing the developer. It then samples cases from production conversation logs, bug reports and support tickets, handwritten sets of 5 to 10 examples, and cases inferred from code. The tool generates a page so the developer can review each input item.
For scoring, the company said developers should use the cheapest viable programmatic comparison first. Only when answers are open-ended should another model be used as a judge, and that judge should not be the same model being tested.
Anthropic’s four conditions for a strong evaluation
Anthropic listed four conditions for a good evaluation:
- The questions should reflect production use.
- Stronger models and higher reasoning intensity should produce higher scores.
- Even the strongest model at the highest setting should still score meaningfully below 100%.
- Run-to-run variation should stay small.
The post also warned that if developers only select questions the model got wrong today, they may end up measuring the model’s current weaknesses rather than the capabilities the application actually needs.

