Anthropic adds build-eval and hillclimb to Claude Code, cutting support-case costs to one-fifth in internal test

Anthropic adds build-eval and hillclimb to Claude Code, cutting support-case costs to one-fifth in internal test

N
News Editor
2026-09-30 11:03:19
Anthropic said on Sept. 28 that it has added two new commands to the claude-api skill available in Claude Code: /claude-api build-eval, which helps developers create evaluation sets for their AI applications, and /claude-api hillclimb, which tunes prompts, models, and parameters based on evaluation results while trying to avoid overfitting to the test set. In an internal customer support example shared by the company, Claude moved a workflow from Opus 4.8 with high reasoning intensity to Sonnet 5 with low reasoning intensity. On previously unseen test tickets, accuracy rose from 78.6% to 90.5%, while cost fell to about one-fifth of the original level. Anthropic said the evaluation covered 44 tickets, with 30 used for tuning and 14 held out for testing. The post also described how hillclimb makes one change at a time and rolls back edits if gains do not carry over to the test set. Anthropic outlined how build-eval samples cases from production conversations, bug reports, support tickets, handwritten examples, and code-derived scenarios, and it listed four traits of a strong evaluation set, including that even the strongest model should still score meaningfully below 100% at the highest reasoning setting.

Anthropic said in a developer blog post on Sept. 28 that it has added two commands to the claude-api skill available in Claude Code: /claude-api build-eval, which helps developers create evaluation sets for their own AI applications, and /claude-api hillclimb, which adjusts prompts, models, and parameters step by step based on evaluation results while trying to prevent overfitting.

In an internal customer support case published by Anthropic, Claude changed a setup that originally used Opus 4.8 with high reasoning intensity into a cheaper configuration based on Sonnet 5 with low reasoning intensity. On unseen test tickets during the tuning process, accuracy increased from 78.6% to 90.5%, while cost dropped to about one-fifth of the starting level.

Support evaluation moved from Opus 4.8 to Sonnet 5

The support evaluation included 44 tickets in total. Of those, 30 were used for tuning and 14 were held out for testing. The starting point was Opus 4.8 at its default high reasoning intensity. On the tuning tickets, decision accuracy was 74.4%, and token cost was 4.6 cents per ticket.

Claude first reviewed the prompt and removed a mandatory tool-calling flow, draft steps, and conflicting rules. It then switched to Opus 5.5 with low reasoning intensity. Accuracy rose to 87.8%, and cost fell to 1.9 cents per ticket.

Anthropic said part of the savings came from Opus 5.5 pricing. Input and output tokens were 20% cheaper than Opus 4.8, while cache reads were 60% cheaper.

After Opus 5.5 met the target, Claude tested the cheaper Sonnet 5 at low reasoning intensity. Accuracy came in at about 88.9%, and cost was cut in half again to about 1 cent per ticket. After adding routing rules and a cross-check on refund limits, Sonnet 5 reached 98.9% on the tuning tickets.

hillclimb changes one thing at a time

Before hillclimb starts, Anthropic said Claude asks the developer what should be optimized, such as improving performance or reducing cost without changing performance. It then randomly splits the evaluation set into a tuning group and a test group.

In each round, Claude reads only the failed records in the tuning group and proposes one modification. If the tuning score improves but the test score does not, the system treats that as a sign of possible overfitting and rolls the change back. A change is kept only when both groups improve.

The post gave an example involving image text recognition. If the evaluation questions happen to require OCR, the tuning process may add an OCR tool to the application, lifting the score on the test but not necessarily helping in real use. Anthropic said failed content should not be pasted directly into prompts, and answers should be structured so the model cannot directly access them.

What happens when scores stall

If scores stop improving for two to three rounds in a row, Claude checks the remaining failed cases one by one and groups them by cause. The goal is to identify whether the issue comes from ambiguous questions, a faulty grader, or problems in the evaluation environment.

Anthropic said it used the same process to tune the claude-api skill itself, raising pass rates from 66% to about 88%. During that work, the company found that one grader required something that did not match the question, while another grader included instructions that conflicted with the official documentation.

How build-eval creates an evaluation set

Anthropic said build-eval starts by interviewing the developer. It then samples cases from production conversation logs, bug reports and support tickets, handwritten sets of 5 to 10 examples, and cases inferred from code. The tool generates a page so the developer can review each input item.

For scoring, the company said developers should use the cheapest viable programmatic comparison first. Only when answers are open-ended should another model be used as a judge, and that judge should not be the same model being tested.

Anthropic’s four conditions for a strong evaluation

Anthropic listed four conditions for a good evaluation:

  • The questions should reflect production use.
  • Stronger models and higher reasoning intensity should produce higher scores.
  • Even the strongest model at the highest setting should still score meaningfully below 100%.
  • Run-to-run variation should stay small.

The post also warned that if developers only select questions the model got wrong today, they may end up measuring the model’s current weaknesses rather than the capabilities the application actually needs.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
100

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.