TechNewsReel
Live

Cactus Releases Needle 2: A 14MB Agentic LLM for Microcontrollers

The 45-million parameter model enables local tool calling and device control on hardware as small as the ESP32-S3.

TechNewsReel Newsroom · August 10, 2026

Cactus has released Needle 2, an open agentic large language model designed to bring tool-calling capabilities to extremely resource-constrained hardware. The release marks a shift toward "edge of the edge" computing, allowing AI agents to operate locally on wearables and microcontrollers without requiring cloud connectivity.

The model features 45 million parameters and is distributed as a single 14MB binary, compressed to CQ2-bit using Cactus Quants. This footprint allows the model to function on budget hardware, including newer microcontrollers such as the ESP32-S3, while operating within 28MB of session RAM.

Technical Foundation and Performance

Needle 2 is built upon findings from the Simple Attention Network (arXiv:2607.18363), focusing specifically on structured extraction, device use, and tool calling. This specialization allows it to maintain high throughput across various low-power platforms.

Performance data provided by Cactus shows the model reaching 500 tokens per second on a Raspberry Pi 5. On budget smartphones, such as the Samsung A-Series, speeds range between 300 and 700 tokens per second. The highest performance is seen on VR devices, with the Meta Quest 3S and Apple Vision Pro delivering between 400 and 1,500 tokens per second.

The Shift to Local Agency

Most agentic LLMs—models capable of interacting with interfaces and executing tools—require significant memory and compute, typically forcing a reliance on high-end GPUs or cloud APIs. By shrinking the requirements to 28MB of session RAM, Cactus is moving agentic AI beyond high-end smartphones and into the realm of embedded systems.

This transition is critical for the development of private, low-latency device control. When an agent can process a command and trigger a tool locally on a microcontroller, it eliminates the latency and privacy risks associated with sending sensor data to a remote server. This has immediate implications for smart home automation and wearable AI, where immediate response times are essential for usability.

Future Outlook

As the successor to the original Cactus Needle, Needle 2 establishes a baseline for how small a functional agent can be. While the developer claims the model competes with larger models like FunctionGemma 270M and Apple FM on certain benchmarks, these specific comparisons have not been independently verified.

Industry observers will now be watching to see how developers integrate Needle 2 into real-world robotics and IoT ecosystems. The primary challenge remains whether a 45M-parameter model can maintain sufficient reliability across complex, multi-step tool chains in diverse environments.

Sources

Get a notification when a big story breaks. A few a day at most — no spam.