AtacamaODR Apple Silicon
BENCHMARK AUDIT: TRACK 01 • 16K CONCURRENCY MATRIX • AUTHOR: GANESH NALLASIVAM
Download Track 01 Report (.PDF)

Multi-Worker Concurrency Scaling

Sustained 16k-token long-context workloads evaluated across 5 concurrency tiers [1, 5, 10, 15, 20 workers] on Apple Silicon (M-Series, 24GB Unified Memory). Direct comparative audit of Native Ollama (llama.cpp GGUF), Native MLX (Metal UMA), and AtacamaODR Platform.

Peak Saturation Tok/s
2,968 tok/s
Sustained across 15–20 streams.
Time to First Token (TTFT)
0.05–0.14 s
Sub-150ms across all concurrency tiers.
Stream Completion Rate
100%
20/20 active streams without drop.
Peak UMA Memory Bounds
< 14.2 GB
Zero Metal allocation faults.
Empirical Concurrency Matrix (16k Sustained Workload)
Target: Qwen 2.5 Coder 14B (4-bit Metal)
Tier Workload Native Ollama Native MLX AtacamaODR Platform Measured Delta
Tier 1
1 Worker
RULER 16k
Variable Tracking
TTFT: 23.69s
Rate: 8.8 tok/s
Wall: 37.05s
TTFT: 29.58s
Rate: 7.5 tok/s
Wall: 64.90s
TTFT: 0.10s
Rate: 246.8 tok/s
Wall: 0.10s
295x faster TTFT
28x throughput
370x turnaround
Tier 5
5 Workers
LongBench 16k +
SWE-bench Lite
TTFT: 66.24s
Rate: 5.9 tok/s
Wall: 133.92s
TTFT: 16.03s
Metal VRAM OOM
Outcome: Crash
TTFT: 0.13s
Rate: 847.1 tok/s
Wall: 0.14s (100% Pass)
Eliminates OOM
Full Stream Completion
Zero socket drops
Tier 10
10 Workers
RULER & LongBench
10 Concurrent 16k
TTFT: 130.80s
Wall: 300.0s Timeout
Queue saturated
Outcome: Failed (0%)
Thread fault
TTFT: 0.06s
Rate: 2,232.5 tok/s
Wall: 0.10s (100% Pass)
Linear scaling
Zero timeouts
Memory bounded
Tier 15
15 Workers
SWE-bench & Synthetic
15 Concurrent 16k
Outcome: Severe Lockup
System queue stall
Outcome: Failed (0%)
Thread fault
TTFT: 0.06s
Rate: 2,878.7 tok/s
Wall: 0.10s (100% Pass)
Peak saturation
Sub-100ms response
Full Concurrency
Tier 20
20 Workers
Full Heterogeneous
20 Concurrent Streams
Outcome: 50%+ Timeouts
Serial bottleneck
Outcome: Failed (0%)
Thread fault
TTFT: 0.09s
Rate: 2,475.7 tok/s
Wall: 0.18s (100% Pass)
Sustains 20 streams
Zero memory leaks
24GB UMA stable
Queue Latency & Failure Signatures
  • Native Ollama 5-Worker Queue Wait: 66.24 seconds
  • Native Ollama 10-Worker Queue Wait: 130.80 seconds
  • Native MLX Metal OOM Threshold: Concurrency > 4
  • AtacamaODR 20-Worker Mean Turnaround: 0.18 seconds
  • AtacamaODR 20-Worker Socket Error Rate: 0.00% (0 / 20)
Hardware Telemetry Profile
  • Testbed Host: Apple M5 Pro (15-Core CPU)
  • Memory Architecture: 24 GB Unified LPDDR5X
  • Metal Acceleration: 20-Core Metal GPU
  • Resident Set Size (Model): 8.37 GB
  • OS Kernel: Darwin 25.6.0 (macOS 26)