Cactus Needle 3 is an 8-29MB model family that does not chat. It handles tool calls, structured JSON extraction, and text embeddings on microcontrollers, wearables, and automotive systems. The entire model is a single binary with no runtime dependencies. You pick a depth (2 to 20 layers) and get a corresponding model size and capability level.
The trade is explicit: Needle sacrifices general reasoning and conversational fluency to match DeepSeek V4 Flash on automation benchmarks while running in a fraction of the memory footprint. This is not a scaled-down GPT. It is a purpose-built inference engine for constrained environments where latency, cost, and offline operation matter more than open-ended dialogue.
Why Small Models for Tool Calling Matter
Frontier models carry alignment tax, safety layers, and billions of parameters optimized for chat. When your workload is “parse a voice command into two function calls” or “extract invoice fields from a photo,” you pay for capacity you do not use.
Needle inverts the priority: tool call accuracy and structured output compliance come first. The model does not guess. If no tool matches the input, it returns an empty list. If the schema requires a date field, the decoder grammar guarantees a parseable date or fails early.
This matters for edge deployment:
- Latency: No network round trip. A smart home command executes in milliseconds.
- Cost: No per-token API charges. The model runs on the device.
- Privacy: Voice commands and sensor data never leave the hardware.
- Reliability: Offline operation survives network outages and API rate limits.
The engineering question is how you compress a model to 8MB without losing the ability to handle multi-step tool orchestration and structured extraction.
Architecture: Simple Attention Network and Depth Slicing
Needle is built on a Simple Attention Network (SAN), a variant that prunes layers and attention heads aggressively while preserving the ability to map input tokens to function signatures and JSON schemas.
Depth as a Tunable Parameter
One set of weights supports every depth from 2 to 20 layers. You choose the layer count at inference time, not training time. A 2-layer model fits in 8MB and handles single-step tool calls. A 20-layer model uses 29MB and handles multi-step orchestration with context carryover.
This is not quantization alone. The model is trained to degrade gracefully as you remove layers. Each depth level is a distinct model with its own accuracy profile.
| Depth (Layers) | Size (MB) | Use Case | Latency (ms) |
|---|---|---|---|
| 2-4 | 8-12 | Single tool call, classification | <10 |
| 6-10 | 14-20 | Multi-tool sequences, extraction | 15-30 |
| 12-20 | 22-29 | Complex orchestration, embeddings | 40-80 |
Constrained Decoding for Structured Output
Needle enforces output schemas at decode time. You declare a JSON shape (fields, types, constraints) and the model generates tokens that satisfy the grammar. If the schema requires an ISO date, the decoder only considers tokens that form valid dates.
This is not post-processing. The model does not generate free text and then parse it. The grammar is part of the inference loop. Invalid outputs are impossible, not just unlikely.
For tool calls, the schema is the function signature. The model sees the available functions, their argument types, and the user input. It returns a list of calls with typed arguments or an empty list if no function matches.
Compression Techniques and What Gets Pruned
Getting to 8MB requires more than quantization. Needle removes entire capability classes that automation workloads do not need.
What Gets Removed
- Conversational context tracking: No multi-turn dialogue state. Each input is independent.
- Open-ended generation: No creative writing, summarization, or explanation. Output is always structured.
- Safety alignment layers: No refusal logic, no content filtering. The model assumes the tool set is the security boundary.
- Multilingual embeddings: English-only for tool calls. Extraction supports additional languages but with reduced accuracy.
What Stays
- Function signature matching: The model maps natural language to function names and argument slots.
- Type coercion: Strings to dates, numbers, enums. The model infers types from context.
- Multi-step sequencing: “Dim the bedroom and lock the door” becomes two calls in order.
- Negative matching: “Clean the kitchen but not the bedroom” generates a call with exclusion parameters.
The pruning is task-specific. If your workload is tool calling and extraction, you get frontier-class accuracy. If you need reasoning, explanation, or chat, you get nothing.
Deployment Shape and Integration
Needle ships as a single binary. No Python runtime, no ONNX conversion, no dependency hell. You link it into your application and call inference functions directly.
Example Integration: Smart Home Command
use needle::{Model, ToolCall};
fn main() {
let model = Model::load("needle-3-8mb.bin").unwrap();
let tools = vec![
Tool::new("dim_lights", vec!["room", "level"]),
Tool::new("lock_door", vec!["door_id"]),
];
let input = "dim the bedroom to 30% and lock the front door";
let calls: Vec<ToolCall> = model.infer(input, &tools).unwrap();
for call in calls {
println!("{}: {:?}", call.function, call.arguments);
}
}
Output:
dim_lights: {"room": "bedroom", "level": 30}
lock_door: {"door_id": "front"}
The model does not return explanations, confidence scores, or alternative interpretations. It returns executable calls or an empty list.
Failure Modes and Escalation
Small models fail differently than frontier models. They do not hallucinate tool calls. They return empty lists when uncertain. But they also miss nuance and context that larger models catch.
Common Failure Patterns
- Ambiguous references: “Turn off the lights” when multiple rooms have lights. The model picks the first match or returns an error.
- Implicit arguments: “Set the thermostat” without a temperature. The model omits the argument or uses a default if the schema allows it.
- Complex conditionals: “If the door is unlocked, lock it.” The model does not evaluate conditionals. It generates a lock call unconditionally.
Escalation Strategy
Needle is not a general-purpose assistant. It is a first-pass filter. When it returns an empty list or a low-confidence call, you escalate to a larger model or prompt the user for clarification.
A typical flow:
- Needle attempts local inference (8-29MB, <50ms).
- If the output is empty or ambiguous, escalate to a cloud model (GPT-4, Claude).
- If the cloud model also fails, fall back to a UI prompt.
This keeps 90% of requests local and fast while handling edge cases with heavier infrastructure.
Fine-Tuning for Custom Tool Sets
Needle 3 supports fine-tuning on the Cactus platform for $19 per run. You provide examples of user inputs and expected tool calls. The platform generates training data, runs the fine-tune, and returns a custom model binary.
Fine-tuning is useful when:
- Your tool set has domain-specific terminology (medical devices, industrial equipment).
- You need to handle abbreviations or jargon that the base model misses.
- You want to enforce specific argument formats (internal IDs, custom enums).
The fine-tune does not add layers or increase model size. It adjusts weights to prioritize your tool signatures over the base distribution.
Observability and Debugging
Needle does not expose token probabilities, attention maps, or intermediate activations. You get a list of tool calls or an error. This is intentional: the model is an inference engine, not a research artifact.
For debugging, you log inputs and outputs and compare them to expected behavior. If the model misses a tool call, you add that example to your fine-tuning dataset. If it generates an invalid argument, you tighten the schema constraints.
There is no prompt engineering. You do not tune system messages or few-shot examples. You define the tool set and let the model map inputs to calls.
Performance Benchmarks
Cactus claims Needle 3 matches DeepSeek V4 Flash on automation tasks. The comparison is narrow: tool call accuracy and structured extraction, not general reasoning or chat quality.
Key metrics:
- Tool call accuracy: 94% on single-step calls, 89% on multi-step sequences (6-layer model).
- Extraction F1 score: 0.91 on invoice fields, 0.88 on booking data (10-layer model).
- Latency: 15ms median for single tool call on a Raspberry Pi 4.
These numbers are task-specific. Needle does not compete with frontier models on MMLU, HumanEval, or conversational benchmarks. It competes on the subset of tasks that matter for edge automation.
Security Boundaries
Needle has no safety alignment. It does not refuse requests, filter outputs, or detect adversarial inputs. The security boundary is the tool set you expose.
If you give Needle access to a “delete_all_files” function, it will call that function when the input matches. The model does not evaluate consequences or ask for confirmation.
This is a feature for constrained environments. You define the action space at integration time, not inference time. The model is a router, not a gatekeeper.
For production deployments, you layer security at the tool execution boundary:
- Allowlists: Only specific tool calls are executable.
- Rate limits: Throttle calls per user or device.
- Audit logs: Record every tool call for forensic analysis.
Technical Verdict
Use Needle 3 when:
- You need tool calling or structured extraction on edge devices (wearables, automotive, smart home).
- Latency and offline operation matter more than conversational quality.
- Your tool set is well-defined and changes infrequently.
- You can tolerate occasional misses and escalate to a larger model when needed.
Avoid Needle 3 when:
- You need open-ended reasoning, explanation, or multi-turn dialogue.
- Your tool set is large (>50 functions) or changes frequently.
- You require safety alignment, content filtering, or refusal logic.
- You need multilingual support beyond English.
Needle is not a replacement for frontier models. It is a specialized inference engine for automation workloads where size, speed, and cost constraints dominate. If your deployment fits that profile, Needle delivers frontier-class accuracy at a fraction of the resource cost. If your workload requires general intelligence, you still need a large model.