An 8GB Jetson Orin Nano looks like a tight target for modern local LLMs. In practice, it can stretch much further than the model-file size alone suggests. In our standardized benchmark pass on Turing Pi 2.5, all seven Q4_K_M models we tested loaded with full logged GPU layer placement at a 4,096-token context setting, including Qwen3 8B. Several of the smaller models also completed real filled prompts far beyond 4k, with four reaching the 32k stage three times.
But “does this model fit in 8GB?” is only the starting question. Context length can add gigabytes of K/V cache, runtime buffers consume their own memory, and every service running on the Jetson shares the same LPDDR5 pool. A model that looks comfortable at 4k can hit a wall at 16k or 32k, while architectures with different attention and state layouts can behave very differently.
So we tested the combinations that actually matter in a deployment. We benchmarked all seven models under identical 4k conditions, filled their context windows with tokenizer-counted prompts, mapped how context growth changed the memory budget and recorded the failure or safety-stop points we actually reached, compared full and partial GPU offload on Qwen3 8B, and ran identical GGUF files through both Ollama and llama.cpp. The result is a practical map of what fits, how fast it runs, how much filled context we could reproduce, and where an 8GB Jetson starts running out of room.
1. The short answer: all seven Q4 models ran at 4k
For the main comparison, we used the same clean system state and the same llama.cpp settings for every model: 4,096 configured context, full GPU offload request, FP16 K/V cache, six CPU threads, batch 512, microbatch 512, Flash Attention enabled, temperature 0, and seed 42.
The pp512 and tg128 columns come from five-repetition llama-bench runs. They are the cleanest apples-to-apples throughput numbers in this article. Request-level timings are separate measurements because chat templates and tokenization produced different effective prompt lengths. These controlled response runs used short prompts at a 4,096-token context setting; the filled-context tests come later.
| Model | GGUF size | GPU layers | pp512 | tg128 | Mean TTFT | Lowest MemAvailable |
| Qwen3 1.7B Q4_K_M | 1.11 GB | 29/29 | 1,408.3 ± 46.5 tok/s | 39.42 ± 0.10 tok/s | 0.62 s | 3,680 MiB |
| Gemma 4 E2B Q4_K_M | 3.46 GB | 36/36 | 926.6 ± 28.6 tok/s | 26.67 ± 0.13 tok/s | 1.48 s | 1,829 MiB |
| Llama 3.2 3B Instruct Q4_K_M | 2.02 GB | 29/29 | 813.7 ± 12.6 tok/s | 25.24 ± 0.01 tok/s | 0.73 s | 2,875 MiB |
| Ministral 3 3B Instruct Q4_K_M | 2.15 GB | 27/27 | 766.5 ± 6.4 tok/s | 23.27 ± 0.07 tok/s | 1.51 s | 2,863 MiB |
| Qwen3 4B custom Instruct Q4_K_M | 2.50 GB | 37/37 | 609.7 ± 8.7 tok/s | 19.53 ± 0.02 tok/s | 1.26 s | 2,913 MiB |
| Phi-4-mini-instruct Q4_K_M | 2.49 GB | 33/33 | 705.1 ± 8.4 tok/s | 19.43 ± 0.03 tok/s | 0.88 s | 2,853 MiB |
| Qwen3 8B Q4_K_M | 5.03 GB | 37/37 | 355.5 ± 3.2 tok/s | 12.64 ± 0.01 tok/s | 2.03 s | 860 MiB |

Figure 1. Full-GPU generation throughput from five-repetition llama-bench tg128 tests. All seven models were measured on the same clean node and software stack.
The result that matters most for capacity is the last row. The 5,027,784,224-byte Qwen3 8B GGUF loaded with all 37 logged layers on GPU and completed all three controlled response runs. Its 4k configuration used a 4,455.34 MiB CUDA model buffer, a 576 MiB GPU K/V allocation, and a 100.01 MiB GPU compute buffer.
That does not mean 8B is the obvious choice. Qwen3 1.7B generated more than three times as fast in the controlled microbenchmark and left much more system memory available. The table measures capacity and throughput, not answer quality.
2. System state is part of the memory budget
For the final benchmark pass, we removed nonessential resident AI workloads before testing. GNOME/Xorg and normal system daemons remained active, so this was still a realistic Linux inference node rather than a stripped benchmark image. At baseline, Linux reported 6,068,916 KiB of MemAvailable and zero swap in use.
We standardized that state because an earlier run with additional AI workloads resident produced different capacity results for some of the same models. On an 8GB unified-memory Jetson, model capacity is also system capacity.
That makes resident workload state part of the benchmark configuration. Model weights, K/V cache, runtime buffers, the operating system, and other applications all compete for the same physical memory pool.
3. Why GGUF file size is not the memory requirement
The Orin Nano 8GB uses shared LPDDR5 memory. CPU tasks, GPU inference, Linux, the graphical session, filesystem cache, inference buffers, and other applications all draw from the same physical pool.
A useful first approximation is:
Required memory ≈ model weights
+ K/V cache or model state
+ runtime and compute buffers
+ operating system and other processes
Quantization reduces the size of the weights, but it does not remove the rest of the equation.
3.1 Context can consume gigabytes by itself
For a conventional full-attention transformer, an approximate FP16 K/V-cache budget for one sequence is:
K/V bytes per token = 2 × layers × K/V heads × head dimension × bytes per element
For Qwen3 8B, with 36 transformer layers, eight K/V heads, head dimension 128, and FP16 cache elements:
2 × 36 × 8 × 128 × 2 bytes = 147,456 bytes/token = 144 KiB/token
4,096 context = 576 MiB
8,192 context = 1,152 MiB
16,384 context = 2,304 MiB
32,768 context = 4,608 MiB
The first two values match the K/V allocations we observed for Qwen3 8B in the final benchmark. The model itself did not change between those runs. Only the amount of context memory changed.
Not every architecture logs its state in exactly the same way. Gemma 4 E2B, for example, reported a much smaller set of hybrid state/K/V buffers than the conventional dense-transformer examples. Treat runtime logs as architecture-specific measurements rather than forcing every model into one K/V formula.
3.2 A configured context is not a filled context
Setting -c 32768 proves only that the runtime attempted to reserve a 32k window. It does not prove that the model successfully processed a 32k prompt.
For the context tests below, we used tokenizer-counted text fixtures and recorded the final prompt length after the model’s chat wrapper. We counted a context stage only when the runtime actually processed the long prompt and completed the request.
4. Test system and methodology
This article continues the setup established in our Jetson Orin Nano Super installation guide, Jetson AI-node preparation guide, and Ollama deployment guide. Those guides cover installation and service setup. Here we focus on capacity and performance.
| Component | Final benchmark configuration |
| Carrier | Turing Pi 2.5 |
| Compute module | NVIDIA Jetson Orin Nano 8GB |
| OS | Ubuntu 24.04.5 LTS |
| JetPack / Jetson Linux | JetPack 7.2.1 / Jetson Linux R39.2.1 |
| Kernel | 6.8.12-1021-tegra |
| CUDA | 13.2 |
| Power profile | 25W, mode 1 |
| llama.cpp | build 10706, commit 1a07bfa5f, CUDA backend |
| Ollama | 0.34.0 |
| Clean baseline | 6,068,916 KiB MemAvailable, 0 KiB swap used |
| Desktop | GNOME/Xorg remained active |
| GPU temperature | about 55.8 °C at baseline; highest recorded GPU/TJ temperature 85.9 °C |
The highest recorded temperature is included for transparency, but this was not a sustained thermal-throttling study. We did not lock CPU/GPU clocks or run a thermal soak.
4.1 Controlled 4k benchmark
Every primary model used:
- 4,096 configured context
-ngl 999to request full GPU placement- FP16 K/V cache
- six CPU threads
- batch 512 and microbatch 512
- Flash Attention enabled
- temperature 0 and seed 42
- the same benchmark prompt content
- a target of 128 generated tokens
- five
llama-benchrepetitions forpp512andtg128 - three independent cold load-and-response attempts
A cold run means a fresh runtime/model process with no model already resident. We did not drop Linux filesystem caches between runs.
Swap was also not reset between cells. Desktop and system pages accumulated in swap during the ordered benchmark, so peak swap figures are useful as system-state observations but not as a fair per-model memory ranking. That is why the main comparison emphasizes logged CUDA allocations and MemAvailable instead.
4.2 Exact model artifacts
All seven primary files were Q4_K_M. We pinned the exact files rather than relying on model display names alone.
| Model | Exact GGUF | Source revision | SHA-256 prefix |
| Qwen3 1.7B | Qwen3-1.7B-Q4_K_M.gguf | Qwen/Qwen3-1.7B-GGUF@7fb011e9 | 228fb5627f75... |
| Llama 3.2 3B Instruct | Llama-3.2-3B-Instruct-Q4_K_M.gguf | bartowski/Llama-3.2-3B-Instruct-GGUF@5ab33fa9 | 6c1a2b411610... |
| Ministral 3 3B Instruct | mistralai_Ministral-3-3B-Instruct-2512-Q4_K_M.gguf | bartowski/mistralai_Ministral-3-3B-Instruct-2512-GGUF@0a903530 | fec9d28c7f8d... |
| Gemma 4 E2B | google_gemma-4-E2B-it-Q4_K_M.gguf | bartowski/google_gemma-4-E2B-it-GGUF@81012ba3 | 923c4c86177d... |
| Phi-4-mini-instruct | microsoft_Phi-4-mini-instruct-Q4_K_M.gguf | bartowski/microsoft_Phi-4-mini-instruct-GGUF@7ff82c2a | 01999f17c39c... |
| Qwen3 4B custom Instruct | Qwen3-4B-Q4_K_M.gguf | Qwen/Qwen3-4B-GGUF@bc640142 | 7485fe6f11af... |
| Qwen3 8B | Qwen_Qwen3-8B-Q4_K_M.gguf | bartowski/Qwen_Qwen3-8B-GGUF@0b69f75b | 54fffa050078... |
Similar model names do not guarantee identical weights, conversions, tokenizers, or runtime behavior. The abbreviated checksums are included here for readability; the benchmark archive retains the full SHA-256 values.
5. How much filled context actually worked?
The 4k benchmark gives us a fair cross-model baseline. The context sweep answers a different question: how far can each exact model go when the prompt really fills the window?
For each primary model that passed 4k, we progressed through larger tokenizer-counted inputs. The important number below is the maximum reproducibly completed in this run, not a claimed architectural maximum.
| Model | Maximum repeated stage | Runtime prompt tokens | Prefill at maximum | Decode at maximum | GPU K/V or state at maximum | Lowest MemAvailable |
| Qwen3 1.7B | 32k, 3/3 | 32,585 | 787.2 tok/s | 13.21 tok/s | 3,584 MiB | 1,403 MiB |
| Llama 3.2 3B | 32k, 3/3 | 32,609 | 438.2 tok/s | 11.23 tok/s | 3,584 MiB | 486 MiB |
| Ministral 3 3B | 32k, 3/3 | 32,575 | 437.1 tok/s | 11.20 tok/s | 3,328 MiB | 635 MiB |
| Gemma 4 E2B | 32k, 3/3 | 32,591 | 594.0 tok/s | 21.70 tok/s | 204 MiB logged hybrid state/K/V | 2,259 MiB |
| Phi-4-mini | 16k, 3/3 | 16,195 | 504.1 tok/s | 12.55 tok/s | 2,048 MiB | 1,435 MiB |
| Qwen3 4B | 16k, 3/3 | 16,201 | 436.0 tok/s | 12.03 tok/s | 2,304 MiB | 1,326 MiB |
| Qwen3 8B | 8k, 3/3 | 8,010 | 320.5 tok/s | 10.50 tok/s | 1,152 MiB | 336 MiB |

Figure 2. Maximum context stage completed three times with a real filled prompt. Qwen3 8B stopped at 8k because the next stage was blocked by the predefined memory-safety threshold, not because 16k was proven impossible.
5.1 Where 32k hit the memory boundary
Phi-4-mini and Qwen3 4B both completed filled 16k prompts three times, then failed twice at 32k during model/context initialization.
Phi-4-mini attempted to allocate a 4,096 MiB FP16 CUDA K/V buffer. Qwen3 4B attempted a 4,608 MiB K/V allocation. Both returned cudaMalloc failed: out of memory, followed by failed to allocate buffer for kv cache.
Those are useful capacity failures because the requested context allocation itself explains the pressure. They do not require a speculative fixed allocator ceiling.
Qwen3 8B is different. It completed filled 8k prompts three times, but MemAvailable fell to 336 to 409 MiB. Our predefined safety rule prevented a 16k attempt once available memory dropped below 512 MiB. Therefore 8k is the tested maximum for this run, not a proven model or runtime limit.
The 32k passes also need context. Three successful requests show reproducibility at that stage, not long-term production reliability. By the 32k stages, roughly 0.9 to 1.1 GiB of swap was in use. Because swap was not reset between cells, treat that as system-state context rather than per-model swap consumption. We did not perform an hours-long soak test.
6. Qwen3 8B: full GPU or partial offload?
Partial CPU/GPU offload can make a larger model easier to place when GPU-side allocations are tight. The trade-off is performance.
We tested the same Qwen3 8B Q4_K_M file in two configurations at 4,096 context:
| Configuration | GPU layers | Request decode | TTFT | CUDA model buffer | GPU K/V | CPU K/V |
| Full offload | 37/37 | 12.23 tok/s | 2.03 s | 4,455.34 MiB | 576 MiB | 0 MiB logged |
| Partial offload | 24/37 | 5.85 tok/s | 3.02 s | 3,015.58 MiB | 368 MiB | 208 MiB |

Figure 3. Request-level generation throughput for the same Qwen3 8B Q4_K_M file. Both configurations completed three responses.
Partial offload cut the CUDA model buffer by about 1.41 GiB, but generation fell to less than half the full-offload rate in these requests. Prompt processing also dropped from 288.5 tok/s to 200.2 tok/s, while TTFT increased from 2.03 to 3.02 seconds.
That makes partial offload a useful fallback, not the preferred configuration here. Once the node had enough clean memory, full 37/37 GPU placement both worked and performed substantially better.
7. Ollama versus llama.cpp with identical GGUF files
Runtime comparisons are easy to get wrong when two tools silently use different model files. For this test we used the same GGUF checksum in both runtimes for two models:
- Qwen3 1.7B: SHA-256
228fb5627f7510b8b3516cdb6435e4b0d2a2bf330fe5b0ab19284a3570a8bb1f - Qwen3 4B: SHA-256
7485fe6f11af29433bc51cab58009521f205840f5b4ae3a32fa7f92e8534fdf5
Both engines received a raw 512-token prompt, a 128-token target, context 4096, temperature 0, seed 42, six threads, batch 512, and a full GPU placement request.
| Model | Runtime | Cold model load | Cold decode | Warm mean TTFT | Warm mean decode |
| Qwen3 1.7B | Ollama 0.34.0 | 23.27 s | 33.68 tok/s | 0.052 s | 34.16 tok/s |
| Qwen3 1.7B | llama.cpp 1a07bfa5f | 3.01 s | 35.93 tok/s | 0.037 s | 36.64 tok/s |
| Qwen3 4B | Ollama 0.34.0 | 39.52 s | 16.35 tok/s | 0.077 s | 16.33 tok/s |
| Qwen3 4B | llama.cpp 1a07bfa5f | 5.52 s | 18.65 tok/s | 0.064 s | 18.72 tok/s |

Figure 4. Warm request decode for identical GGUF files under matched request settings. These are measured cases, not a universal runtime ranking.
llama.cpp loaded both test files faster and produced higher warm decode throughput in this setup. That is an observation about these specific builds and loading paths, not proof that llama.cpp is always faster than Ollama.
There is another measurement wrinkle: both runtimes reused 511 of 512 prompt tokens during the warm repeats. Warm prompt-evaluation figures therefore describe prompt-cache behavior and should not be presented as full 512-token prefill throughput. The controlled llama-bench pp512 values in Section 1 are the better cross-model prefill comparison.
For deployment, the trade-off is broader than speed. Ollama provides a convenient long-running model service and API, while llama.cpp exposes more direct control over low-level inference settings. Our Ollama on Jetson guide covers that service setup separately.
8. What this means on Turing Pi 2.5
The Turing Pi carrier does not add memory to the Jetson or make its GPU faster. Each compute module keeps its own CPU, accelerator, and memory.
What the board changes is the deployment layout. A Jetson can be dedicated to GPU inference while other nodes handle application logic, databases, retrieval, automation, monitoring, or other services. That matters because every service left on the Jetson competes directly with model weights, K/V cache, and runtime buffers for the same shared memory.
We did not benchmark relocating those services to a companion node and measure the before/after latency or memory savings, so this article does not claim a quantified multi-node benefit. The point is architectural: a multi-node system gives you the option to keep the Jetson focused on inference when that separation is useful.
If a standalone Jetson with enough free memory already handles the complete application, there is no capacity advantage to adding another node. The multi-node benefit appears when separating workloads is useful for the system you are actually building.
9. Choosing a configuration from these results
There is no single winner because we did not evaluate answer quality. The useful decision is to match the measured capacity and latency to the job.
| If your priority is… | What the measurements show |
| Highest tested generation throughput | Qwen3 1.7B led the controlled set at 39.42 tok/s tg128 and completed filled 32k prompts 3/3. |
| More memory headroom at 4k | Qwen3 1.7B left the most measured MemAvailable of the primary models. Smaller models generally gave the system more room for context and other services. |
| The largest tested Q4 model with full GPU placement | Qwen3 8B completed 4k 3/3 with 37/37 logged GPU layers at 12.64 tok/s tg128. |
| Long filled prompts | Qwen3 1.7B, Llama 3.2 3B, Ministral 3 3B, and Gemma 4 E2B each completed their 32k stage 3/3. |
| A 4B-class model with longer context | Qwen3 4B completed 16k 3/3; its 32k K/V allocation failed twice. |
| A partial-offload fallback | Qwen3 8B at 24/37 GPU layers worked 3/3 but request decode fell to 5.85 tok/s. |
| Running alongside substantial companion services | Leave headroom for the operating system, runtime buffers, caches, and every service that shares the Jetson’s memory. |
Do not read this table as a quality ranking. We did not score reasoning, coding, factual accuracy, instruction following, or long-context comprehension. A smaller model that is faster and easier to host may still be the wrong model for a particular task, and a larger model may not justify its memory cost if the smaller one already meets the application’s quality target.
10. Troubleshooting model-load failures
CUDA reports out of memory while Linux still shows some available memory
Record the exact failed allocation, MemAvailable, active inference processes, context setting, and runtime logs. CUDA may be unable to satisfy a large model or K/V allocation even while Linux still reports some memory available for the rest of the system.
Before concluding that a model cannot fit, retry under a documented baseline and compare the requested allocation with the memory already consumed by the runtime and resident services.
The model loads at 4k but fails at a larger context
Look at the K/V allocation first. Phi-4-mini and Qwen3 4B both loaded at 16k, then failed at 32k when their FP16 K/V allocations grew to 4,096 MiB and 4,608 MiB respectively.
Reducing context can free a large amount of memory without changing the GGUF file at all.
A model name works in one runtime but fails in another
Verify the exact file checksum. Display names are not enough to prove that two tools are loading identical weights or conversions.
Then compare context length, K/V precision, GPU placement, runtime version, and resident workload state before blaming the runtime itself.
Partial offload works but feels much slower
That is expected when more layers execute on the CPU and splitting execution between CPU and GPU adds overhead. Measure it. In our Qwen3 8B test, partial offload completed every request but generated at less than half the full-GPU request-level rate.
11. Limits of this benchmark
This dataset answers a specific capacity question rather than every deployment question.
- We did not perform an answer-quality ranking.
- Three repeated requests at key stages do not establish long-term production reliability.
- We did not test multi-user concurrency.
- We did not run a sustained thermal soak or controlled throttling experiment. The highest recorded GPU/TJ temperature was 85.9 °C.
- Qwen3 8B above 8k was not tested because the predefined memory-safety cutoff stopped the sweep.
- Optional Q3, Q8, Gemma 4 E4B, and Ministral 3 8B boundary artifacts were not included in the final standardized pass.
- We did not benchmark actually relocating services to a companion Turing Pi node.
12. What we learned
The benchmark gives a practical answer to the original question.
All seven primary Q4_K_M models, including Qwen3 8B, loaded with full logged GPU layer placement and completed three response tests with a 4,096-token configured context on this Jetson Orin Nano 8GB.
Context was the next boundary. Qwen3 1.7B, Llama 3.2 3B, Ministral 3 3B, and Gemma 4 E2B completed filled 32k stages three times. Phi-4-mini and Qwen3 4B reached 16k reproducibly before 32k K/V allocations failed. Qwen3 8B reached 8k reproducibly before the safety cutoff prevented a larger test.
The useful rule is not “an 8GB Jetson supports models up to X billion parameters.” The model file, K/V cache, runtime buffers, operating system, and every other service all share the same memory pool. Treat the complete deployment as the unit you are sizing.
Related articles
- NVIDIA Jetson Orin Nano Super on Turing Pi 2.5: Complete Setup Guide
- Preparing NVIDIA Jetson as an AI Node on Turing Pi 2.5
- Run Ollama on Jetson Orin Nano: Build a Local AI Server on Turing Pi 2.5
- Local AI on Turing Pi with NVIDIA Jetson: When Edge AI Makes Sense
FAQ
Can an 8GB Jetson Orin Nano run an 8B LLM?
Yes, for the exact configuration tested here. Qwen3 8B Q4_K_M loaded with 37/37 logged GPU layers at 4,096 context and completed three controlled responses. It also completed filled 8k prompts three times. We did not test 16k because available memory fell below the predefined safety threshold after the 8k stage.
That result applies to this GGUF, runtime, power profile, and clean system state. It is not a guarantee that every 8B model or every 8GB Jetson deployment will behave the same way.
Is the GGUF file size the amount of memory the model needs?
No. The file represents the quantized model weights. Inference also needs K/V or model-state memory, runtime and compute buffers, plus memory for Linux and every other process on the machine.
Which tested model was fastest?
Qwen3 1.7B produced the highest controlled llama-bench tg128 result at 39.42 tok/s. That is a speed result, not a quality ranking.
Which models completed a real 32k prompt?
Qwen3 1.7B, Llama 3.2 3B, Ministral 3 3B, and Gemma 4 E2B each completed the final 32k filled-context stage three times. Phi-4-mini and Qwen3 4B failed at 32k during K/V allocation after completing 16k. Qwen3 8B was stopped at 8k by the safety threshold, so its behavior above 8k remains untested here.
Does partial GPU offload help an 8B model?
It can reduce GPU-side memory use. Qwen3 8B at 24/37 GPU layers completed all three requests, but request decode fell from 12.23 tok/s with full offload to 5.85 tok/s with partial offload. On this clean node, full offload was the better-performing configuration.
Is Ollama slower than llama.cpp on Jetson?
In the two identical-GGUF cases measured here, llama.cpp loaded the model faster and had somewhat higher warm generation throughput. The runtimes use different engine revisions and loading paths, so these measurements should not be generalized into a universal ranking.
Does adding another Turing Pi node increase the Jetson’s model memory?
No. The nodes have independent memory. Another node can host application services, databases, retrieval, automation, or other workloads so they do not consume the Jetson’s shared memory, but it does not combine its RAM with the Jetson’s GPU memory pool.