Speed · measured on our machine, published with the method

Fast, because there is no trip.

Cloud AI has to receive your request before it can think about it. Atlas starts on the chip already in front of you. Here is what that looks like, where we win, where we do not yet, and exactly how we measured it.

0.35 sto the first wordqwen3.5:4b, warm, median of 10
17 sto cut and grade 2 min 20 s of 1080pAuto Cinematic, about 8× real time
0.19 sto auto-enhance an 8 MP photofull-resolution JPEG, median of 5
0 sspent uploadingyour files are already on your disk

Measured September 29, 2026 on one Apple M5 Pro MacBook (15-core CPU, 16-core GPU, 24 GB), on AC power, Ollama 0.34.3, everyday apps left open.

Chat

As fast as the fastest cloud model. Without sending a word.

For a short answer, Atlas and the quickest cloud models start typing at about the same moment. The difference is where your words had to go to get there. Press run: the bars fill at a quarter of real speed so you can see them.

0 ms

Studio

Where local pulls away.

Photos and footage are heavy. A cloud editor has to receive every byte before it starts. Atlas reads them straight from your disk and does the work on your chip.

Auto Cinematic

17.1 s

Two minutes and twenty seconds of 1080p footage, analyzed, cut, graded and rendered into a 30 second 2.39:1 film. About 8× faster than real time.

Photo Auto Enhance

0.19 s

An 8 MP iPhone photo, fully corrected and saved at full resolution. A 36.6 MP photo takes 0.56 s.

Instant visuals

0.03 ms

To pick the right built-in simulation for a question, with no model call. A custom diagram shows its first node in 1.7 s.

Editing video?

The cloud has to receive it first.

Before a cloud editor can touch your footage, every byte has to travel up your internet connection. Set your file size and your upload speed. This is arithmetic, not a benchmark.

A minute of 4K iPhone video is roughly 170 to 400 MB depending on settings.

Cloud editor27 minjust to upload, before editing starts
Atlas0 sthe file is already here. Editing starts now.

Why

Where the time goes, and where it doesn't.

Cloud AI waits on

  1. Distance. Your request crosses the internet to a data center and the answer crosses it back.
  2. A shared queue. The same GPUs serve millions of people. At busy hours you wait your turn.
  3. Uploads. Photos, video and documents have to arrive before work can begin.
  4. Metering. Every token costs the provider money, so usage is capped and throttled.

Atlas skips it because

  1. The model is already here. Barx keeps it warm in your memory between questions.
  2. Nobody else is in line. Your chip works for you alone.
  3. The file is already here. Studio reads straight from your disk.
  4. The hardware is matched. Barx fits the model, context and GPU layers to your machine every time.

Honest trade-off: the largest cloud models are bigger than anything a laptop can hold, and on hard reasoning they can be stronger. Atlas is built for speed, privacy and everyday work, and hands off to the web when an answer needs fresh facts.

Benchmarks

Every number, with its conditions.

Median of repeated runs unless noted. One machine, one date. Smaller machines run smaller models, which Barx chooses automatically, so your numbers will differ. Run them yourself and tell us what you get.

Local chat

Measureqwen3.5:4bqwen3.5:9bRuns
First word, warm0.345 s0.696 s10 each
Output speed57.7 tokens/s34.8 tokens/s10 each
A natural answer (65 to 132 tokens), total1.92 s2.83 s10 each
Exactly 150 tokens, total3.42 s5.47 s10 each
First word from cold, including model load2.66 s3.54 s3 each

Atlas's own system prompt (about 459 tokens), thinking off, context 8192, 15 threads, all layers on the GPU, as Barx chose for this machine. Cold means unloaded from Ollama with files still in the macOS cache, not a fresh boot.

Visuals

MeasureMedianRangeRuns
Pick a built-in visual (no model call)0.026 ms0.019 to 0.045 ms12 questions × 2,000
Custom diagram: first node ready1.68 s1.49 to 1.90 s10
Custom diagram: complete, explicit request3.27 s2.54 to 4.31 s10

Engine time only; drawing on screen was not measured. 11 of 12 test questions matched a built-in. 12 of 14 diagram attempts returned a diagram.

Auto Cinematic video

ClipInputOutputTotalFaster than real time
Barx demo140 s, 1080p3030 s, 1080p, 2.39:1, graded17.1 s8.2×
GlassBox explainer112.8 s, 1080p3030 s, 1080p, 2.39:1, graded17.4 s6.5×
Atlas teaser89.2 s, 720p3030 s, 720p, 2.39:1, graded9.5 s9.4×

Uncached analysis plus full render with app defaults (Teal & Orange, grain on), median of 3. HDR iPhone footage was not in this set.

Photo Auto Enhance

PhotoSizeTime
iPhone JPEG8.3 MP0.194 s
iPhone JPEG9.1 MP0.192 s
iPhone HEIC7.2 MP0.344 s
iPhone HEIC18.4 MP0.619 s
JPEG36.6 MP0.562 s

Full-resolution JPEG output, median of 5.

Image generation

RunSizeTime
Cold, including model load1024 × 102448.2 s
Warm1024 × 102436.8 s

Z-Image Turbo, 4-bit, one run each. Our slowest engine today, and the one we are working on.

Reaching a cloud AI (network only)

EndpointDNSTCPTLSConnected
api.openai.com2.8 ms20.0 ms31.3 ms54.9 ms
api.anthropic.com2.6 ms22.5 ms32.8 ms58.9 ms
generativelanguage.googleapis.com3.1 ms26.4 ms38.7 ms70.3 ms

Fresh connection each time over 5 GHz Wi-Fi, median of 7. No keys or data were sent. This is time before any model starts working.

Reproduce it

# Chat: first word and speed, with Ollama installed
curl -s http://127.0.0.1:11434/api/pull -d '{"model":"qwen3.5:4b"}'
curl -s http://127.0.0.1:11434/api/chat -d '{
  "model": "qwen3.5:4b", "think": false, "stream": false,
  "options": {"num_ctx": 8192},
  "messages": [{"role": "user", "content": "Explain photosynthesis in three sentences."}]
}' | python3 -c 'import json,sys; d=json.load(sys.stdin); print(d["eval_count"]/d["eval_duration"]*1e9, "tokens/s")'

# Network floor to a cloud AI (no key needed; the server refuses you)
curl -so /dev/null -w 'dns %{time_namelookup}  tcp %{time_connect}  tls %{time_appconnect}\n' https://api.openai.com/v1/models

# Studio benchmarks run inside Atlas. We will publish the full scripts with the public release.