Cactus Compute, a Y Combinator S25 on-device AI startup, has introduced Needle 3, a model designed for tool calling and structured extraction rather than general chat. In the example cited by BlockBeats, when a user says "set the living room light to 30%," the model identifies the right function and fills in parameters such as "living room" and "30%." The report said Needle 3 resembles the recently popular Jev in some ways, but follows a different path: Jev scores candidate decisions directly, while Needle 3 still generates token by token, with outputs constrained to tool calls that software can execute directly. One set of Needle 3 weights can be trimmed from 2 to 20 layers. The 2-layer version is about 9 MB, while the full 20-layer model is about 35 MB, allowing deployment choices across devices including watches, Raspberry Pi boards, and smartphones. After separate fine-tuning for Android operation commands, the 4-layer, 29 million-parameter version posted a self-tested score of 62.5, above DeepSeek V4 Flash at 60.5.
Cactus Compute, a Y Combinator S25 startup focused on on-device AI, has released Needle 3, a model aimed at tool calling and structured extraction instead of general chat.
According to BlockBeats, the model is built to turn a user request into executable parameters. In the example provided, if a user says, "set the living room light to 30%," Needle 3 identifies the relevant function and fills in parameters such as "living room" and "30%."
The report said Needle 3 shares some similarities with Jev, which has recently drawn attention, but the two take different approaches. Jev directly scores candidate decisions. Needle 3 still generates outputs token by token, but those outputs are restricted to tool calls that a program can execute directly.
Needle 3 also comes with a flexible size range. A single set of weights can be trimmed to anywhere from 2 to 20 layers. The 2-layer version is about 9 MB, while the full 20-layer version is about 35 MB. That allows developers to choose a version based on available compute across devices such as watches, Raspberry Pi systems, and phones.
After separate fine-tuning for Android operation commands, the 4-layer version with 29 million parameters achieved a self-tested score of 62.5, compared with 60.5 for DeepSeek V4 Flash.
This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan. Disclaimer:
The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.
Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.