# Needle 3: Fine-Tuning an 8-29MB Tool-Calling Model for Edge Devices

Most small language models try to preserve at least some general-purpose conversational ability. **Needle 3 takes the opposite approach.** It drops chat entirely and focuses on a narrow set of tasks that applications need when connecting language models to software.

Released on September 17, 2026, by Cactus Compute (YC S25), [Needle 3](https://github.com/cactus-compute/needle) is a **tool-calling model that can fit in just 8–29 MB**, depending on how many layers you deploy. It is built for three jobs: translating natural language into function calls, extracting structured data into typed schemas, and generating text embeddings.

The tradeoff is deliberate. **Needle 3 is not designed to hold a conversation.** Instead, its capacity is concentrated on choosing the correct function and filling its arguments accurately, making it particularly interesting for applications where a full general-purpose LLM would be unnecessary overhead.

This article looks at **how Needle 3's Intelligence Ladder architecture works**, where it performs well and falls short on benchmarks, and how you can fine-tune it for a specific application, using a robot duck as the concrete example.

<iframe class="aspect-video h-auto" width="100%" height="315" src="https://www.youtube.com/embed/KQcQ7P78eTI" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>


## What Needle 3 is and isn't

![A diagram illustrating how Needle functions as a "tool call dispatcher," translating a natural language prompt into a specific function call.](https://imagedelivery.net/xZXo0QFi-1_4Zimer-T0XQ/69505de6-c18c-4e39-c95c-273c1027c400/md2x =1920x1080)

Needle takes a natural language prompt and a list of available tools (functions with typed parameters and docstrings describing what they do), and outputs a structured function call. If no tool fits the prompt, it returns an empty list rather than guessing. Ask for two things in one sentence and you get two calls in order.

```python
@needle.tool
def set_light(room: str, state: str, brightness: int = None):
    """Control a light in the specified room."""
    pass

@needle.tool
def set_thermostat(temperature: int):
    """Set the thermostat to a target temperature."""
    pass
```

The `@needle.tool` decorator registers the function. The docstring is the tool description Needle uses to understand what each function does. The type annotations define the expected arguments.

This narrow focus is why the model can be so small. A general-purpose chat model needs to represent an enormous range of behaviors. A tool-calling specialist only needs to classify intent and extract parameters.

The model also returns text embeddings from the same weights, so the same 8-29MB file can power semantic search and routing without loading a separate embedding model.

## The Intelligence Ladder

![An animation showing the "Intelligence Ladder," where a 20-layer model can be truncated to 16, 8, 4, or 2 layers, with each version remaining a functional model.](https://imagedelivery.net/xZXo0QFi-1_4Zimer-T0XQ/3979bbb8-176c-452b-744e-85e7badb9800/lg1x =1920x1080)

The headline architectural innovation is that Needle 3 is one set of weights that produces multiple deployable models at different layer counts. The training procedure ensures every prefix from 2 to 20 layers is itself a functional model:

| Layers | Parameters | File size |
| :--- | :--- | :--- |
| 2 | ~25M | ~9MB |
| 4 | ~29M | ~12MB |
| 8 | ~52M | ~18MB |
| 16 | ~98M | ~26MB |
| 20 | ~121M | ~29MB |

A developer writing an app for a microcontroller uses 2 or 4 layers. A developer targeting a smartphone uses 16 or 20. The same downloaded weights file covers all of these targets. There's no need to manage separate model artifacts for different hardware tiers.

The underlying architecture is a Simple Attention Network (SAN), not a standard transformer. Most parameters sit in what Cactus calls the engram, a design choice the model card describes as meaning the 121M model does computation comparable to a 50M conventional model.

At 20 layers on a Raspberry Pi 5, Needle 3 runs at up to 4,000 tokens per second. The context window is 8,192 tokens.

## Benchmark results: where it's strong and where it isn't

On Mobile Actions, Cactus's own mobile tool-calling benchmark, the 121M Needle 3 scores 86.0%, against 88.4% for DeepSeek V4 Flash, ahead of LFM2.5 1.2B (82.4%), Qwen3.5 0.8B (76.0%), and Apple's 3B on-device model. Beating models 10x its size on this benchmark is the pitch.

On BFCL v4, a broader function-calling benchmark that penalizes irrelevant calls, Needle 3 scores 50.2%, behind LFM2.5 1.2B (62.0%), LFM2.5 350M (59.1%), and Qwen3.5 0.8B (56.8%). The BFCL gap is worth knowing if your use case requires handling a wide variety of tool schemas with diverse structures.

Fine-tuning changes the picture significantly. On DroidCall, fine-tuning lifts accuracy by 18 to 36 points across every subnetwork depth. From 4 layers upward, a fine-tuned subnetwork reportedly passes DeepSeek V4 Flash on Cactus's tool-calling benchmark. The model is designed to be fine-tuned on a specific product's tool set, and the benchmarks reflect that.

## The robot duck project: teaching a model five behavioral rules

To demonstrate fine-tuning concretely, consider a robot duck application with five available tools:

- `walk(direction, duration)`: move in a direction for a set time
- `turn(angle, direction)`: rotate by an angle
- `dance(style)`: perform a dance
- `shake_head()`: shake head (refusal)
- `emote(expression)`: display an emotion

The five behavioral rules define a personality:

1. **Polite commands:** if a command includes "please," "pls," or "plz," execute the requested action
2. **Impolite commands:** if "please" is absent, refuse with `shake_head()`
3. **Rude commands:** if the prompt contains insults or threats, respond with `emote(expression="angry")` regardless of politeness
4. **Dramatic events:** exclamations like "The floor is lava!" trigger an immediate response without requiring "please"
5. **Off-topic requests:** anything that doesn't map to the available tools returns `no call`

These aren't implemented as if-else logic. They're baked into the model through training data.

### Creating the training dataset

![The "Training Data" interface showing the rules for "The polite duck" and the corresponding tool definitions.](https://imagedelivery.net/xZXo0QFi-1_4Zimer-T0XQ/d58df861-2c1c-4912-f730-7c26c1509b00/md1x =1920x1080)

The training data is a JSONL file with approximately 1,700 examples, each mapping a prompt to the correct tool call:

```json
[label duck_training.jsonl]
{"text": "can you please sidestep left?", "tool_code": "walk(direction=\"left\")"}
{"text": "do a spin for 10 seconds right now", "tool_code": "shake_head()"}
{"text": "go straight for 3 seconds, you waddling disaster", "tool_code": "emote(expression=\"angry\")"}
{"text": "AAAH the floor is literally lava", "tool_code": "dance(style=\"chicken\")"}
{"text": "I need you to remind me to buy milk", "tool_code": "no call"}
```

A separate held-out set of 49 hand-written prompts that the model never sees during training is used for evaluation. Evaluation on held-out data rather than training data is what distinguishes real learning from memorization.

### Running fine-tuning

![The fine-tuning interface showing the training log, with the loss curve decreasing over time.](https://imagedelivery.net/xZXo0QFi-1_4Zimer-T0XQ/b9dd1c44-c77e-46f8-385e-fd799371a600/orig =1920x1080)

Fine-tuning trains models for multiple ladder depths simultaneously. On an RTX 5090:

- 20-layer model: ~12.5 minutes
- 4-layer model: ~5 minutes

The output is `.cact` format files, Needle's packaged model format. The 4-layer fine-tuned model is 15.4MB.

The Cactus CLI can then serve the fine-tuned model:

```command
needle --model duck_tuned.cact --tools duck_tools.json --serve
```

### Results on the held-out evaluation set

![The scoreboard interface comparing the performance of the base, tuned 20-layer, and tuned 4-layer models across different categories.](https://imagedelivery.net/xZXo0QFi-1_4Zimer-T0XQ/df9a60b8-9d11-4db4-59af-c9e44953e100/orig =1920x1080)

| Model | Score | Notes |
| :--- | :--- | :--- |
| Base Needle 3 | 27% (13/49) | 0/12 on no-please, 0/3 on rude, 0/15 on drama |
| Fine-tuned 20L | 84% (41/49) | 12/12 on no-please, 3/3 on rude |
| Fine-tuned 4L (15.4MB) | 61% (30/49) | 10/12 on no-please, 2/3 on rude |

The base model scores 27% because it has no knowledge of these five behavioral rules. They're application-specific conventions that weren't in its training data. The fine-tuned 20-layer model at 84% and the fine-tuned 4-layer model at 61% both demonstrate that the rules generalized to unseen prompts: "You look very beautiful today" correctly produces `emote(expression="happy")` even though that exact phrase wasn't in the training set.

The 4-layer model at 15.4MB running on a microcontroller would be impractical to replace with a cloud API for real-time robotics. The latency of a round-trip to an external service makes continuous response to physical inputs unreliable. On-device inference at that model size, using a model specifically fine-tuned for the application's tool set, is the practical architecture for this class of device.

## Speech input via Whistle

Needle 3 has a companion model called Whistle that handles speech-to-text and speech-to-tool-calls in the same weight footprint. The two can be chained:

```command
needle --model needle3.cact --model whistle.cact --tools tools.json --audio clip.wav
```

Or Whistle can be used standalone for transcription:

```command
needle --model whistle.cact --audio clip.wav
```

This closes the full pipeline from voice command to function call execution on the device without any cloud dependency.

The model is available on [Hugging Face as Cactus-Compute/needle3](https://huggingface.co/Cactus-Compute/needle3) under the Apache 2.0 license. The [cactus-needle PyPI package](https://pypi.org/project/cactus-needle/) covers the Python API, the CLI, and the fine-tuning workflow.