This Jetson Orin Nano Dev Kit vs Turing Pi 2.5 comparison starts with one important constraint: moving the Jetson between carriers does not give the module a faster GPU, more memory, or a higher nominal compute ceiling. The CPU, GPU, Tensor Cores, and 8GB of LPDDR5 live on the Jetson module itself.
What changes is the system around it.
NVIDIA’s developer kit is designed as a complete single-node development platform. It gives the Jetson direct camera connectors, a 40-pin expansion header, four USB-A ports, DisplayPort, wireless connectivity, and a simple 19V power supply. Turing Pi 2.5 trades some of that bench-friendly I/O for four compute-module slots, per-node NVMe, an onboard managed Ethernet switch, centralized ATX power, and a BMC that can control nodes remotely.
That makes the obvious question surprisingly useful: if the exact same Jetson module runs the exact same workload on both carriers, does anything measurable change?
We ran the same 8GB Jetson Orin Nano module and 500GB NVMe for this Jetson Orin Nano Dev Kit vs Turing Pi 2.5 comparison, on Turing Pi 2.5 and on NVIDIA’s reference carrier, and repeated the same tests on both. The goal is not to prove that one carrier is universally better. It is to separate the things that stay the same from the things that genuinely change when a Jetson moves from a standalone developer kit into a multi-node Turing Pi system.
1. The result at a glance
The comparison uses one physical Jetson module, not two nominally identical samples. That removes module-to-module silicon variation from the most important measurements.
| Measurement | NVIDIA developer kit | Turing Pi 2.5 | Difference |
| llama.cpp prefill, pp512 | 810.57 tok/s median (810.35 to 810.67) | 810.15 tok/s median (809.84 to 810.64) | +0.05% |
| llama.cpp decode, tg128 | 25.276 tok/s median (25.200 to 25.378) | 25.215 tok/s median (25.125 to 25.288) | +0.24% |
| 30-minute sustained decode (median of 360 reps) | 25.304 tok/s; first-to-last 60 reps +0.425% | 25.319 tok/s; first-to-last 60 reps +0.145% | -0.06% |
Peak module temperature (tj) | 78.09 °C (idle 50.22 °C, rise +27.9 °C) | 86.75 °C (idle 54.75 °C, rise +32.0 °C) | -8.66 °C |
Jetson module idle input (VDD_IN) | 4.757 W median | 4.447 W median | +0.31 W (+7.0%) |
Jetson module input under LLM load (VDD_IN) | 21.133 W median | 20.972 W median | +0.16 W (+0.8%) |
| Cold model load, 3B GGUF | 9.681 s median (9.441 to 9.694, n=5) | 9.730 s median (9.508 to 23.597, n=3) | -0.5% |
| NVMe sequential read (1M, fixed protocol) | 2,570 MiB/s (PCIe 3.0 x4) | 2,580 MiB/s (PCIe 3.0 x4) | -0.4% |
| NVMe sequential write (1M, 60 s) | 1,202 MiB/s (PCIe 3.0 x4) | 2,714 MiB/s (PCIe 3.0 x4) | -55.7% (drive-state dependent, see section 6) |
| GbE TCP throughput to the RK1 peer (different path, see section 7) | 934.1 Mbps median (0 retransmits) | 939.0 Mbps median (0 retransmits) | Different path, not directly comparable |
Inference performance was effectively identical. Prefill differed by 0.05% and decode by 0.24%, both inside the run-to-run spread, and neither carrier lost throughput over the 30-minute sustained run. Median cold model loading, sequential reads, 4K random I/O, and GbE throughput also landed within a few percent of each other (one of three Turing Pi cold loads was a 23.6 s outlier, covered in section 6).
Two results were not close. The dev kit ran the module 8.7 °C cooler at peak, and its 60-second sequential write speed was 56% lower than on Turing Pi. The write gap turned out to depend on drive state and timing rather than on a fixed carrier limit, and sections 6 and 8 explain what we can and cannot say.
The important limit is that this is an n=1 controlled deployment comparison, not a universal benchmark of every Orin Nano carrier. If a performance gap appears, the first suspects are power delivery, cooling, clocks, software state, and background services rather than the carrier somehow changing the Jetson’s underlying compute resources.
2. Tested configuration
Our earlier Jetson articles used the same Orin Nano 8GB development module on Turing Pi 2.5. This comparison reuses that hardware so the new measurements connect directly to the rest of the series.
| Component | Configuration |
| Jetson module | Orin Nano 8GB development module, P3767-0005 |
| Turing Pi | Turing Pi 2.5, board revision 2.5.2 as reported by the BMC, module in Node 2 |
| NVIDIA carrier | Jetson Orin Nano Super Developer Kit reference carrier |
| Storage | Crucial P310 500GB NVMe (CT500P310SSD8, firmware VACR001), same physical drive on both carriers |
| Jetson Linux | R39.2.1 (kernel 6.8.12-1021-tegra, TNSPEC 3767-300-0005-W.1-1-1-jetson-orin-nano-devkit-super-) |
| JetPack | 7.2.1-b49 |
| CUDA | 13.2.2 |
| Power mode | 25W (nvpmodel mode 1: 6x Cortex-A78 to 1344 MHz, GPU to 918 MHz, EMC to 3199 MHz; jetson_clocks not applied) |
| LLM runtime | llama.cpp build 10706, commit 1a07bfa5f, CUDA backend, 6 threads, Flash Attention on |
| Benchmark model | Llama-3.2-3B-Instruct-Q4_K_M.gguf, sha256 6c1a2b41161032677be168d354123594c0e6e67d2b9227c84f296ad037c728ff |
| Turing Pi BMC firmware | 2024.05.1 (daemon 2.3.4, Buildroot 2024.05.1) |
| Developer-kit PSU | Included Lite-On PA-1450-26 adapter (NVIDIA-labeled), 19 V, 2.37 A, 45 W output, 100 to 240 V AC input |
| Turing Pi PSU | Lapcare LPS450 ATX supply (450 W class, 24-pin), 230 V AC input; label lists +12 V 24 A, +5 V 19 A, +3.3 V 16 A max |
| Cooling | Same module heatsink and fan assembly on both carriers, not removed. The fan is driven by nvfancontrol (profile quiet, close-loop) through the module fan header on both carriers. |
| Enclosure | Turing Pi: in its case, with BMC-controlled case fans (not visible to the Jetson). Developer kit: open-air on a desk, no enclosure. |
| Dev-kit wireless | Wi-Fi radio off and disconnected during measurements; the Bluetooth service stop was cancelled by the system and it stayed active during the main measurement windows (the Turing Pi carrier has neither radio). |
Turing Pi figures come from the existing installation: the module was not removed, re-flashed, or reconfigured. Resident workloads and desktop services were stopped for the measurement windows and restarted afterwards, on both carriers. Idle RAM use was not identical, though: 678 MiB median on Turing Pi against 1,508 MiB on the developer kit, so the resident software was close but not the same. The same stop list (ollama, docker, containerd, gdm, gnome-remote-desktop, and others) was applied on both carriers and confirmed inactive before the measurement windows, so those services do not explain the gap: dev-kit memory use was already about 1.5 GiB before they were stopped and stayed there afterwards. Several developer-kit-only units were not on the stop list, including snapd, and the Bluetooth stop was cancelled. We did not save per-process memory for the developer kit run, so we cannot say which software accounts for the difference. Under sustained load, RAM use was nearly the same (3,295 vs 3,265 MiB), so the idle gap reflects background system software rather than the benchmark. Keep that in mind when reading the idle power comparison in section 8.
Before each benchmark pass, we recorded the software state rather than assuming the two boots were identical:
cat /etc/nv_tegra_release
uname -a
sudo nvpmodel -q
sudo jetson_clocks --show
systemctl --failed
systemctl list-units --type=service --state=running
cat /proc/cmdline
sudo nvbootctrl dump-slots-info
sudo nvme smart-log /dev/nvme0
sudo lspci -vv -s <NVME_PCI_ADDRESS>
systemd-analyze
free -h
For GPU, memory, thermal, and module-power logging:
sudo tegrastats --interval 5000
The same NVMe booted unchanged on the developer kit
The physical NVMe is the same in both tests, but that does not guarantee a Jetson installation can move between carriers without intervention.
Our Turing Pi setup used NVIDIA’s Linux_for_Tegra tools with the carrier EEPROM check disabled for the third-party carrier. The TNSPEC above shows the install was flashed with the dev-kit-super configuration, so the boot chain and device tree already matched NVIDIA’s reference carrier.
In practice it booted unchanged. We moved the module and drive to the developer kit and the existing installation came up with the same IP address, the same 25W power mode, the same model file hash, the NVMe at the same PCIe address, and no failed services. No reflash or configuration change was needed, and the kernel log shows two Phy link never came up messages for one PCIe host bridge at boot with no effect on the NVMe or Ethernet links.
That is a result specific to an install flashed with the dev-kit-super configuration. If your Turing Pi install used a different configuration, cloning the NVMe before the move and being prepared to reflash is still sensible.
3. Jetson Orin Nano Dev Kit vs Turing Pi 2.5: why inference should be close
The Jetson Orin Nano is a system-on-module. Its CPU, Ampere GPU, Tensor Cores, and LPDDR5 memory are on the module rather than on the carrier board.
NVIDIA also makes an important point about the current “Super” configuration: the higher-performance Orin Nano Super experience was enabled on existing developer-kit hardware through software. NVIDIA lists up to 67 TOPS for the Super Developer Kit and exposes a 25W power mode for the Orin Nano series. The carrier still has to deliver the required power and support the relevant interfaces, but the performance uplift is not a different GPU soldered onto a different carrier.
For this comparison, both deployments use the same 25W configuration:
sudo nvpmodel -q
The expected result is therefore not “Turing Pi makes the Jetson faster.” The expected result is roughly equivalent compute throughput when the module, power mode, software, cooling, and workload are controlled. The measurements in section 5 match that expectation.
That expectation still needed measurement. A carrier can affect sustained performance indirectly if the module is cooled differently, receives unstable power, runs different background services, or enters a different clock or thermal state.
Storage starts from a similarly matched link
NVIDIA’s developer-kit carrier exposes its primary 2280 NVMe slot over PCIe 3.0 x4. Turing Pi’s current I/O documentation maps Jetson Orin Nano NVMe to PCIe 3.0 x4 as well.
That does not guarantee identical SSD results. Firmware, link negotiation, thermal behavior, filesystem state, and I/O contention can still matter. It does mean that one carrier did not start with an obvious lane-width advantage in this test.
4. Measuring the carrier instead of the benchmark noise
A same-module comparison is only useful if the rest of the test is controlled tightly enough that a small difference means something.
We used the same physical Jetson and NVMe, then kept these variables fixed wherever possible:
- same Jetson Linux and JetPack release
- same 25W
nvpmodelprofile - same llama.cpp build
- same GGUF file and model hash
- same prompt and generated-token counts
- same CPU-thread setting
- same GPU offload configuration
- same set of resident services stopped for the measurement windows (the dev kit’s Bluetooth service could not be stopped and stayed active)
- same module heatsink and fan assembly
- same NVMe filesystem and mount options
- same room and a similar time of day
Not everything could be matched. The Turing Pi sat in its case while the developer kit ran open-air, so the thermal test compares the real deployments, not just the PCB beneath the module. That is still useful, but it must be labelled correctly.
The same rule applies to power. The developer kit uses its included 19V supply, while Turing Pi uses ATX power, so power-conversion losses differ. We did not measure them. We report module input power (VDD_IN) from tegrastats, which is read from the same on-module sensor on both carriers.
5. LLM throughput: does the carrier change inference speed?
What LLMs Actually Fit on a Jetson Orin Nano 8GB? Models, Context, and Runtimes Tested established the Orin Nano 8GB LLM benchmark methodology used here. This comparison reuses the same model, runtime build, pp512/tg128 metrics, and six-thread CUDA setup, but uses a shorter carrier-comparison protocol instead of repeating the full benchmark suite. The protocol is not identical to that benchmark: here we used three llama-bench repetitions and -ngl 99, which is already more than enough for this 29-layer model. Both carriers received identical settings.
Earlier exploratory Turing Pi runs, including one on 2026-09-28, used a less controlled protocol and are not used in the tables. Every Turing Pi figure in this article comes from the fixed-protocol pass run immediately before the module was moved to the developer kit, and the raw logs are retained. The only exception is the earlier sequential read results in section 6, which are labeled as such.
For this carrier comparison, one stable model was enough to expose any meaningful sustained-performance difference while keeping the test short and repeatable.
Both carriers used the same GGUF and llama.cpp binary:
./llama-bench \
-m /path/to/model.gguf \
-p 512 \
-n 128 \
-ngl 99 \
-t 6 \
-fa on \
-r 3
After the system reached a stable temperature, we recorded the median prefill and decode results.
| Carrier | pp512 | tg128 | Run 1 | Run 2 | Run 3 |
| NVIDIA developer kit | 810.573 | 25.276 | 810.353 ± 19.260 / 25.200 ± 0.028 | 810.573 ± 16.665 / 25.378 ± 0.054 | 810.672 ± 16.646 / 25.276 ± 0.049 |
| Turing Pi 2.5 | 810.148 | 25.215 | 810.148 ± 17.51 / 25.125 ± 0.050 | 809.836 ± 18.03 / 25.215 ± 0.046 | 810.638 ± 17.05 / 25.288 ± 0.051 |
Values are pp512 / tg128 in tok/s, with llama-bench’s own standard deviation across its three repetitions. Each run started with the module at idle temperature.
- developer kit pp512: 810.353, 810.573, 810.672, median 810.573; tg128: 25.200, 25.378, 25.276, median 25.276
- Turing Pi pp512: 810.148, 809.836, 810.638, median 810.148; tg128: 25.125, 25.215, 25.288, median 25.215
- same command with llama-bench defaults for everything else (
fa auto): developer kit 812.266 / 25.379, Turing Pi 810.414 / 25.152
The medians differ by 0.05% for prefill and 0.24% for decode, and the run ranges overlap. For reference, What LLMs Actually Fit on a Jetson Orin Nano 8GB? Models, Context, and Runtimes Tested measured the same model on this same node on Turing Pi at 813.7 ± 12.6 pp512 and 25.24 ± 0.01 tg128, so both carriers sit inside normal run-to-run variance.
We therefore do not treat either tiny percentage gap as a carrier advantage; both sit inside the observed run-to-run spread.
We also logged tegrastats during the runs so any meaningful throughput difference could be checked against clocks, thermals, memory pressure, and module power:
sudo tegrastats --interval 5000 | tee tegrastats-carrier.log
Those logs let us compare GPU clocks, CPU clocks, temperatures, memory pressure, and reported module power. A sustained clock difference would be more informative than a single tokens-per-second number.
Sustained load matters more than one short pass
A short benchmark can hide a cooling problem. We ran a fixed workload for 30 minutes and compared the beginning and end of the test.
The useful question is not only “which carrier was faster at minute one?” It is whether either deployment loses throughput once heat accumulates.
| Carrier | First 60 reps (median tg128) | Last 60 reps (median tg128) | Change | Peak temp (tj) | Throttle observed |
| NVIDIA developer kit | 25.272 | 25.379 | +0.425% | 78.09 °C | No throughput decay and no explicit throttle event |
| Turing Pi 2.5 | 25.281 | 25.317 | +0.145% | 86.75 °C | No throughput decay across the 30-minute run |
Both carriers held or slightly improved their decode speed over the run, and neither slowed as heat accumulated. The median of all 360 repetitions was 25.304 tok/s on the developer kit and 25.319 tok/s on Turing Pi.
Measurement details, both carriers: six consecutive 5-minute llama-bench blocks (-p 0 -n 128 -ngl 99 -t 6 -fa on -r 60, 360 measured repetitions over about 30 minutes).
- Turing Pi block means: 25.272, 25.319, 25.314, 25.307, 25.354, 25.320 tok/s. GPU at 99% median utilisation, module
VDD_IN20.972 W median,tjpeak 86.75 °C. - Developer kit block means: 25.245, 25.250, 25.322, 25.305, 25.346, 25.369 tok/s. Module
VDD_IN21.133 W median,tjpeak 78.09 °C. The kernel log recorded ahot-surface-alertcooling-state change (0 to 1 about 80 s into the run, back to 0 as the run ended), with no thermal-throttle event string in the kernel ortegrastatslogs.
This is where a deployment difference is most likely to become visible. The carrier does not change the GPU architecture, but the physical system around the module can change its operating conditions. Here the conditions differed (enclosure and airflow) and the decode speed still did not.
6. NVMe performance: same SSD, same nominal x4 link
The developer kit’s 2280 M.2 slot and Turing Pi’s Orin Nano NVMe path are both documented as PCIe 3.0 x4. We tested the same Crucial P310 so the SSD controller and NAND did not change between runs.
We first verified the negotiated link on each carrier:
sudo lspci -vv | grep -A20 -i 'non-volatile memory'
Measured on Turing Pi 2.5 (from lspci -nn and the PCIe sysfs attributes):
| Device | Current link | Maximum link |
Root port 0004:00:00.0 (NVIDIA 10de:229c) | 8.0 GT/s (PCIe 3.0) x4 | 8.0 GT/s x4 |
NVMe 0004:01:00.0 (Crucial/Micron c0a9:5427) | 8.0 GT/s (PCIe 3.0) x4 | 16.0 GT/s x4 (the drive itself is Gen4-capable) |
Jetson GbE 0008:01:00.0 (Realtek 10ec:8168) | 2.5 GT/s x1 | 2.5 GT/s x1 |
The drive sits on its own root port and negotiated PCIe 3.0 x4 on the Turing Pi carrier, matching the starting expectation. The Gen4 ceiling of the P310 is what the carrier gives up. On the developer kit the NVMe enumerated at the same PCIe address and negotiated the same 8 GT/s x4 link.
We then ran identical fio workloads against a dedicated test file, not the live model file. The asynchronous engine matters: without --ioengine, fio uses the synchronous engine and caps queue depth at 1, so a requested --iodepth=32 never happens. On this drive, 1M sequential read measured 347 MiB/s with the synchronous engine versus 2,558 MiB/s with --ioengine=libaio on our first pass, and 4K random read measured 14,023 versus 104,495 IOPS.
Fixed protocol for both carriers: drop caches before each repetition, idle 30 seconds, run 60 seconds, and log nvme smart-log before and after each block. The sequence was one 1M sequential write (to create the file), five 1M sequential reads, three more 1M sequential writes, three 4K random reads, and three 4K random writes.
For sequential reads:
sudo sync; echo 3 | sudo tee /proc/sys/vm/drop_caches
fio --name=seq-read \
--filename=/srv/benchmark/fio-test.bin \
--size=8G \
--rw=read \
--bs=1M \
--ioengine=libaio \
--iodepth=32 \
--direct=1 \
--runtime=60 \
--time_based \
--group_reporting
We repeated the same protocol with --rw=write, then used --rw=randread and --rw=randwrite with --bs=4k for IOPS, while keeping drive fill state and cooling comparable between carriers.
| Test | NVIDIA developer kit | Turing Pi 2.5 | Difference |
| 1M sequential read (median of 5) | 2,570.2 MiB/s (2,569.6 to 2,570.6), 12.36 ms mean, p99 18.74 ms | 2,579.9 MiB/s (2,575.4 to 2,580.1), 12.32 ms mean, p99 18.22 ms | -0.4% |
| 1M sequential write (median of 4) | 1,202.0 MiB/s (1,197.0 to 1,204.3), 26.43 ms mean, p99 68.68 ms | 2,714.3 MiB/s (2,712.9 to 2,715.6), 11.62 ms mean, p99 19.01 ms | -55.7% |
| 4K random read (median of 3) | 104,518 IOPS (408.3 MiB/s), 299 µs mean | 103,015 IOPS (402.4 MiB/s), 303 µs mean | +1.5% |
| 4K random write (median of 3) | 124,034 IOPS (484.5 MiB/s), 251 µs mean | 119,657 IOPS (467.4 MiB/s), 261 µs mean | +3.7% |
Reads, random reads, and random writes were effectively the same. Sequential write was the exception: all four 60-second write blocks on the developer kit landed near 1,200 MiB/s, while all four on Turing Pi landed near 2,714 MiB/s, and the Turing Pi figure had already reproduced across two separate days (2,724.5 MiB/s on the first pass). Both carriers negotiated the same PCIe 3.0 x4 link, and SMART showed no warnings or errors.
We then ran a follow-up on the developer kit only. Sequential writes started at about 2,716 to 2,727 MiB/s in the first 10 seconds, then dropped to 425 to 525 MiB/s in the middle of the run, which averages out to roughly 1,100 to 1,350 MiB/s over 60 seconds. A 10-minute idle did not bring the early burst back. A read immediately followed by a write did reach 2,720.5 MiB/s. Temperatures stayed between 43 and 67 °C, SMART stayed clean, and the link stayed at 8 GT/s x4 throughout. That pattern is consistent with drive-side write behavior (a fast initial burst followed by a slower sustained rate) and with the state the drive was in, rather than with a fixed carrier limit. We did not record a bandwidth-over-time log on Turing Pi, so we cannot say why its 60-second average stayed near the burst rate. We do not attribute the 56% gap to the carrier. Treat it as a measurement of one drive in one state, not as a developer kit limitation. LLM inference and model loading do not depend on it.
Sequential read is state-sensitive on this drive. On Turing Pi it measured 755 MiB/s on a first clean-state pass after a node power cycle, 1,012 MiB/s on a post-write reread, 1,249 MiB/s after a further 10-minute idle, and 2,558 to 2,580 MiB/s in the fixed-protocol pass and on an earlier first pass. The developer kit landed in the fast state (about 2,570 MiB/s) on all five reads. The spread is consistent with SSD-state sensitivity, but these runs do not isolate the cause, so the fixed protocol above was used on both carriers and the spread is reported rather than treating one read result as a fixed carrier limit.
Model load time
For an LLM server, small SSD differences usually matter most during model loading rather than token generation after the model is resident in memory. We measured cold load time with the page cache dropped:
sudo sync; echo 3 | sudo tee /proc/sys/vm/drop_caches
time ./llama-cli -m /path/to/model.gguf -p "hi" -n 1 -ngl 99 -t 6 -st
| Metric | NVIDIA developer kit | Turing Pi 2.5 |
| Cold model load, 3B GGUF (median) | 9.681 s (9.441 to 9.694, n=5) | 9.730 s (9.508 to 23.597, n=3) |
Model loading was effectively the same, which matches the matched sequential-read results. One of the three Turing Pi runs took 23.6 s with no other change, the same kind of SSD-state sensitivity seen in the read tests. The lower sequential write average on the dev kit does not affect loading a model.
The storage test is still valuable because Turing Pi’s per-node NVMe is one of its core deployment features.
7. Networking: direct GbE versus an onboard managed switch
Both systems expose Gigabit Ethernet, but the topology is different.
The developer kit connects the Jetson through its carrier’s GbE interface. On Turing Pi, every compute node connects through the board’s RTL8370MB-based 1Gbps managed switch, which also supports VLANs and links the nodes to the board’s external Ethernet ports.
The two carriers could not be tested against an identical path. We used the same idle Turing RK1 node as the iperf3 peer for both. On Turing Pi the Jetson reached it through the on-board switch. On the developer kit it reached it over the LAN and the Turing Pi uplink, which adds hops. Treat the comparison as two real deployments, not a measurement of the carriers alone. The same wired router ping is included as a cleaner latency check.
On the peer:
iperf3 -s -p 5201
On the Jetson:
iperf3 -c <RK1_PEER_IP> -t 60
iperf3 -c <RK1_PEER_IP> -u -b 900M -t 60
ping -c 100 -i 0.2 <RK1_PEER_IP>
ping -c 100 -i 0.2 <ROUTER_IP>
Wi-Fi was disabled or disconnected during the test, and the Turing Pi switch was left in a documented configuration.
| Network test | NVIDIA developer kit | Turing Pi 2.5 | Difference |
| TCP throughput to the RK1 peer | 934.1 Mbps median (933.7 to 937.9, 0 retransmits) | 939.0 Mbps median (936.0 / 939.0 / 939.0, 0 retransmits) | Different path, not directly comparable |
| UDP, 900 Mbps target | 848.4 Mbps, 0% loss, 0.006 ms jitter | 900.0 Mbps, 0% loss, 0.009 to 0.016 ms jitter | -5.7% |
| UDP, unthrottled flood | 865.9 Mbps, 0% loss | 924.2 Mbps, 0% loss | -6.3% |
| Median ping to the RK1 peer | 0.274 ms | 0.274 ms | 0.0% (rounding coincidence, see note) |
| Median ping to the router (wired) | 0.797 ms | 0.798 ms | -0.1% |
TCP throughput was within 0.5% on both paths and neither saw a single retransmit. The UDP figures were 6% lower on the developer kit, but that path goes through extra hops and we did not isolate why the achieved rate was lower, so do not read it as a carrier difference. Latency to the router was identical to the millisecond. The identical 0.274 ms RK1 medians are a coincidence of rounding: the Turing Pi raw median was 0.2735 ms, and the means and maxima differed (0.311 vs 0.275 ms mean, 0.599 vs 0.320 ms max).
Turing Pi 2.5 measurement details: the Jetson’s enP8p1s0 is a 1000 Mb/s full-duplex link into the board’s RTL8370MB managed switch; the on-board peer was an idle Turing RK1 node on the same switch with no VLANs configured (the switch’s managed/unmanaged selector is a hardware switch and its state is not readable through the BMC API or the firmware in use). TCP ran 60 s per run, UDP used a documented 900 Mbps target and then an unthrottled flood, and latency is the median of 100 pings, 200 ms apart. The board’s 1 GbE uplink, not the switch, is the ceiling for anything off-board.
Raw GbE speed is only one part of the comparison. Turing Pi’s larger advantage is architectural: the Jetson can communicate with three neighboring compute modules through the same integrated switch, and the switch can be managed and segmented with VLANs without adding a separate switch for the cluster.
The developer kit is simpler when the Jetson is the whole system. Turing Pi becomes more interesting when the Jetson is one service inside a larger system.
8. Module power and thermals
NVIDIA documents 7W to 25W power options for the Orin Nano Super Developer Kit, plus a MAXN SUPER mode. We tested the 25W profile only. A configured 25W mode is a clock and power budget, not a meter reading: the module drew about 21 W median under sustained load on both carriers while the profile allows 25 W.
For each carrier we logged two states with tegrastats:
- ten minutes of settled headless idle
- sustained llama.cpp load after temperatures stabilize
| State | NVIDIA developer kit | Turing Pi 2.5 |
Jetson module input at idle (VDD_IN) | 4.757 W median | 4.447 W median |
Jetson module input under sustained LLM load (VDD_IN) | 21.133 W median | 20.972 W median (mean 20.815 W, max 21.331 W) |
Idle temperature (tj) | 50.22 °C median | 54.75 °C median |
Peak temperature (tj) | 78.09 °C (cpu peak 74.72 °C) | 86.75 °C (cpu peak 83.06 °C) |
| Temperature rise, idle to peak | +27.9 °C | +32.0 °C |
| Fan under sustained load | about 3,900 to 4,060 RPM, PWM 163 to 166 of 255 | about 4,050 RPM, PWM 137 of 255 |
Full detail:
| Rail / sensor | Developer kit idle | Developer kit sustained load | Turing Pi idle | Turing Pi sustained load |
VDD_IN (module input) | 4.757 W median | 21.133 W median | 4.447 W median (4.407 to 4.647) | 20.972 W median |
VDD_CPU_GPU_CV | 0.443 W median | 9.377 W median | 0.440 W median | 9.388 W median |
VDD_SOC | 1.491 W median | 5.429 W median | 1.442 W median | 5.413 W median |
tj / gpu temperature | 50.22 / 50.06 °C median | 76.91 °C median, 78.09 °C peak | 54.75 °C median, 55.28 °C max | 83.5 °C median, 86.75 °C peak |
cpu temperature | 48.78 °C median | 73.59 °C median, 74.72 °C peak | 53.53 °C median, 54.0 °C max | 79.97 °C median, 83.06 °C peak |
| RAM used | 1,508 MiB median | 3,295 MiB median | 678 MiB median | 3,265 MiB median |
Both idle captures were 10 minutes (601 samples on the developer kit, 596 on Turing Pi). The developer kit’s idle window started soon after the preceding benchmark, so its first samples were still cooling (maximum 61.5 °C); the median is unaffected.
The developer kit ran cooler under sustained load: 78.09 °C peak against 86.75 °C, with the fan at about the same RPM. The same heatsink and fan assembly was used on both carriers, so the most likely explanation is the airflow around the board (open-air on a desk against a closed case) rather than the carrier electronics, but we did not test that, and we did not log room temperature. These are on-module sensor readings (VDD_IN, tj, cpu, gpu), so read the 8.7 °C gap as a difference between two real deployments, not as a property of either carrier.
Idle module input was 0.31 W (7%) higher on the developer kit. The idle RAM and resident-service differences in section 2 may account for part of that, and we did not separate them. Under sustained load the difference was 0.16 W (0.8%).
VDD_IN covers the Jetson module only. It does not include the carrier, SSD, fans, other nodes, or PSU losses, and we did not measure whole-system draw. A single-node module reading also says nothing about the fixed overhead of the Turing Pi board itself (BMC, switch, fans), and it should not be misread as the energy cost of adding another Jetson to an already-running Turing Pi.
9. What changes between the Jetson Orin Nano Dev Kit and Turing Pi 2.5
The benchmark numbers answer whether the carrier changes measurable performance. The more important buyer question is what each carrier lets you build around the same module.
| Area | Jetson Orin Nano Super Developer Kit | Turing Pi 2.5 |
| Primary role | Single-node development and prototyping | Multi-node edge AI, robotics, self-hosting, and managed deployment |
| Compute slots | 1 Jetson | 4 module slots |
| Module mix | Jetson only | Supported Jetson, RK1, and CM4 modules can be mixed |
| NVMe | 2280 PCIe 3.0 x4 plus a smaller Key-M slot | One 2260/2280 M.2 position per node; Orin Nano mapped at PCIe 3.0 x4 |
| Ethernet | 1x GbE | Integrated managed 1GbE switch, VLAN support, 2 external GbE ports |
| Remote node power | External methods required | BMC power control per node |
| Remote node reset | External methods required | BMC/API reset control |
| Recovery path | USB-C directly from Jetson carrier to host | BMC selects and routes the Jetson recovery path to Turing Pi USB-C OTG, host still runs NVIDIA tools |
| Camera I/O | 2x 22-pin MIPI CSI | No equivalent direct per-Jetson CSI connectors; USB or network cameras can be used where the chosen slot and system topology fit |
| Expansion | Jetson 40-pin header | 40-pin GPIO on Node 1, additional I²C/GPIO connectors, and 2x Mini PCIe; I/O is slot-specific |
| USB | 4x USB 3.2 Type-A plus USB-C modes | Slot-specific USB topology, including USB on Node 1 and 4x USB 3.0 on Node 4, plus switchable USB 2.0 and Mini PCIe USB paths |
| Display | DisplayPort | HDMI on Node 1 with a Jetson-specific switch configuration, plus DSI; our test Jetson was in Node 2, so display output was not tested in this comparison |
| Wireless | Key-E Wi-Fi/Bluetooth module included | Mini PCIe expansion can add Wi-Fi, Bluetooth, cellular, or other interfaces as needed |
| Power | Included 19V supply | Central 24-pin ATX power |
This is why a port-count comparison alone is misleading. The developer kit concentrates a rich set of interfaces around one Jetson. Turing Pi spreads I/O across a four-node platform, so the slot you choose matters, but it is not limited to headless server use. Node 1 provides an HDMI path with a Jetson-specific switch configuration, USB, and board-level GPIO access, while the wider board adds I²C/GPIO expansion, Mini PCIe, NVMe, managed networking, and other compute nodes around it. We did not test the Node 1 display path because our Jetson was installed in Node 2.
The hard difference is direct Jetson-specific I/O. Turing Pi does not reproduce the developer kit’s two MIPI CSI connectors or give every slot the developer kit’s exact 40-pin and USB layout. If a project depends on those exact interfaces on the same Jetson, NVIDIA’s carrier remains the simpler fit.
Where the developer kit is simpler
For direct CSI-camera bring-up, experiments tied specifically to NVIDIA’s Jetson header layout, or a desk setup where several peripherals need to plug straight into the same Jetson carrier, NVIDIA’s developer kit is easier.
Two MIPI CSI connectors sit directly on the developer kit. The Jetson 40-pin header is immediately available for GPIO and serial buses. Four USB-A ports, DisplayPort, and bundled wireless make one-node prototyping very straightforward.
That convenience is important, but it should not be read as “developer kit for robotics, Turing Pi for servers.” Many robotics and edge systems use USB cameras, network sensors, GPIO/I²C devices, wireless links, local displays, and multiple cooperating services rather than requiring the developer kit’s exact CSI and header layout.
Where Turing Pi 2.5 gains ground
Turing Pi 2.5 is built around multiple independent compute nodes, but it can still expose the kinds of interfaces an edge or robotics system needs. Node 1 adds an HDMI path, USB, and a 40-pin GPIO path; the board also exposes additional I²C/GPIO connections and Mini PCIe expansion, while other slots provide different USB, SATA, and storage options. The HDMI circuit has a Jetson-specific switch configuration in Turing Pi’s documentation, but we did not test display output in this comparison. The practical trade-off is that you plan the slot around the I/O instead of getting every interface on one carrier.
That opens up designs that are awkward on a standalone developer kit. An Orin Nano can own perception, speech, or local AI while a neighboring RK1, CM4, or second Jetson handles control software, databases, APIs, telemetry, logging, automation, storage, or another accelerated workload. The nodes communicate through the integrated switch, each can have its own NVMe, one ATX supply powers the platform, and the BMC can power-cycle or reset nodes remotely.
For a robot, kiosk, vision appliance, local AI box, or managed edge system, that can be a more useful architecture than putting every responsibility on one Jetson. The advantage is not higher CUDA throughput. It is keeping essentially the same Jetson compute performance while giving the rest of the system somewhere to live.
10. What the BMC can actually do with a Jetson
“Remote management” can become vague very quickly, so it is worth separating documented capability from things the BMC does not do.
| Capability | Status | What it actually means |
| Per-node power on/off | Verified on the test Jetson (BMC 2024.05.1, daemon 2.3.4) | The BMC API and tpi tooling can control node power state. Power off stops the Jetson in about 4 s; power on returns it to SSH in about 98 s, booting unchanged from the same NVMe. |
| Per-node reset | Verified on the test Jetson | opt=set&type=reset&node=1 reboots the node; back in about 85 s with a new boot ID. |
| UART | Verified: works for the RK1 nodes, not for the Jetson | The RK1 nodes’ Linux consoles are readable (and writable) through the BMC UART. The Jetson node returns an empty buffer, and a BMC UART write produced nothing on the Jetson’s ttyTCU0, so treat the BMC UART as unusable for Jetson console output on this board and firmware. |
| USB recovery routing | Already proven in our setup | tpi usb flash --node N exposes the Jetson recovery path so an external Linux host can see NVIDIA APX |
| NVIDIA Jetson flashing performed by BMC | No | NVIDIA’s Linux_for_Tegra or SDK tooling still runs on the external host |
| Generic raw-image Flash Node feature | Yes, but different | Turing Pi can write raw .img files to supported node storage; this should not be described as the NVIDIA Jetson flashing workflow |
| Per-node watt telemetry from BMC | Not documented, confirmed absent | BMC power state is not the same as wattage; use Jetson VDD_IN telemetry. Wall draw was not measured. |
| USB host/device modes | Exposed by BMC, partially verified with this Jetson | Device and flash modes can be set and read back (Flash/Device, routed to AlternativePort or Bmc). Putting a node in the USB host role is rejected outright: Selecting one of the nodes as USB Host role is not supported by the current hardware. |
| Node index conventions | Practical warning | power uses 1-based node1..node4 while uart, reset and usb use a 0-based node id; nodeinfo is deprecated and returns zeros. A wrong index silently targets a different node. In our session a power command aimed at the Jetson shut down the neighbouring RK1 instead. |
The recovery distinction matters. In our Turing Pi Jetson setup guide, the BMC put the selected node into the right USB recovery path and the Ubuntu host detected the module as an NVIDIA APX device. The actual Jetson Linux installation was still performed by NVIDIA’s flashing tools on that host.
That workflow is useful because the module can stay physically installed in Turing Pi. It is not the same as the BMC understanding Jetson Linux and flashing the module by itself.
11. Setup and recovery are different experiences
The developer kit is hard to beat for the first hour with a Jetson. It arrives as a complete supported platform with the module, reference carrier, cooler, wireless hardware, and power supply. NVIDIA’s JetPack 7.2 download notes now direct Jetson Orin Nano Developer Kit users to the unified JetPack ISO image on a USB stick rather than the older SD-card-image path.
Turing Pi is more involved because it is not a single-board Jetson appliance. The module has to be installed in a node slot, cooled correctly, connected to the right per-node storage, powered through the platform, and brought through the BMC-controlled recovery path when a low-level reflash is required.
Our earlier setup also found a concrete example of why this matters. On our board, recovery transfers repeatedly failed with the Jetson in Node 1, while moving the same module and NVMe to Node 2 allowed flashing to complete. We explicitly treated that as an observation from one system, not evidence that Node 1 fails universally.
That extra setup is the price of a more flexible managed platform. If all you want is one Jetson on a desk, the developer kit is still the faster path; if the Jetson is becoming part of a larger system, the additional integration buys you centralized management and room to expand.
12. The cost comparison needs the right baseline
The cheapest way to run one Orin Nano is not automatically the same as the best way to deploy several compute nodes.
As of September 27, 2026, NVIDIA’s Jetson FAQ lists the Jetson Orin Nano Super Developer Kit at $399. The same FAQ lists the production Jetson Orin Nano 8GB module at $399 at 1KU+, which is volume suggested pricing rather than a direct consumer-retail quote. Turing Pi currently lists the Turing Pi 2.5 board at $279 before the compute module, PSU, storage, and cooling.
Those prices make one conclusion straightforward: Turing Pi 2.5 is not a cheaper one-node Jetson carrier.
If you already own the developer kit, keeping the Jetson where it is costs nothing. Moving its module to Turing Pi adds a $279 cluster board plus whatever power, cooling, case, and storage components your build still needs.
If you are starting from zero and only need one Jetson, the developer kit packages far more of the required hardware into one purchase.
Turing Pi’s economics change as the system grows because the board, BMC, switch, ATX power path, and enclosure can serve several nodes. There is no honest universal “break-even at two nodes” number, however. The answer depends on which modules you install, what storage each node needs, whether you already own a PSU and case, and how much value you place on centralized management. This article also does not compare PSU efficiency or whole-system power draw.
Price is therefore part of the deployment decision, not a benchmark score.
13. Who should keep the NVIDIA developer kit?
Keep the developer kit when your priority is the shortest path from one Jetson to NVIDIA’s direct carrier I/O.
That is especially true if you need:
- the two MIPI CSI camera connectors directly on the Jetson carrier
- the developer kit’s exact Jetson 40-pin GPIO and peripheral layout
- several USB peripherals connected directly to the same Jetson without planning around slot mapping
- direct DisplayPort rather than Turing Pi’s Node 1 HDMI path
- the simplest supported single-node setup
- a compact bench system with its own included PSU and cooler
- the most straightforward first-time development and recovery workflow
- one Orin Nano with no need for neighboring compute nodes or remote management
Those are meaningful advantages for certain prototypes, especially when direct CSI or NVIDIA’s reference I/O layout is central to the project. They are not a requirement for robotics or edge AI in general. Turing Pi can still support local display, USB peripherals, GPIO/I²C devices, expansion hardware, networking, storage, and Jetson acceleration; it simply distributes those capabilities across the board instead of putting all of them on one carrier.
14. Who should move to or add Turing Pi 2.5?
Turing Pi makes more sense when you want the Jetson to be the accelerator inside a system you intend to keep, manage, and expand rather than the entire system by itself.
The board becomes especially useful when you need:
- two to four independent compute nodes in one Mini-ITX platform
- a mix of Jetson, RK1, and CM4-class nodes
- a Jetson dedicated to AI while other nodes run control software, APIs, storage, monitoring, automation, or orchestration
- per-node NVMe storage
- integrated Gigabit Ethernet between nodes
- VLAN-capable managed switching
- centralized ATX power
- remote power and reset through the BMC
- a recovery path that does not require physically removing the Jetson module
- HDMI, USB, and GPIO access on Node 1 when a local display or direct control interface is useful
- additional I²C/GPIO and Mini PCIe expansion for buttons, displays, wireless, cellular, or other hardware
- slot-specific USB 3.0 connectivity for cameras, storage, capture devices, or other peripherals
- a compact platform for distributed AI, robotics, computer vision, self-hosting, Kubernetes, storage, monitoring, or automation
The slot mapping matters, so a peripheral-heavy build should be planned around the node that exposes the interfaces it needs. One practical caveat from our setup is that recovery transfers repeatedly failed with the Jetson in Node 1 and succeeded after moving the same module and NVMe to Node 2. That was one system, not evidence of a universal Node 1 limitation, but it is worth planning around if your build depends on Node 1’s display and GPIO I/O. In return, Turing Pi can turn one Jetson project into a complete edge system without changing the Jetson’s measured inference performance.
This is the reason our Jetson series uses Turing Pi. The board does not make CUDA faster. It makes the system around CUDA much more capable: storage, networking, management, expansion, and other compute can all live in the same platform.
15. What Turing Pi 2.5 does not change
The move adds system-level capability, not new silicon inside the Jetson. Do not expect:
- more GPU compute from the same Jetson module
- more LPDDR5 memory
- magically higher LLM tokens per second
- a cheaper single-node Jetson build
- the Turing Pi BMC to replace NVIDIA’s Jetson flashing tools
- BMC-level per-node watt telemetry
- the same two direct MIPI CSI connectors as NVIDIA’s carrier
- the developer kit’s exact USB and 40-pin I/O layout on every slot
Those limits are narrower than saying Turing Pi cannot handle displays, robotics, or peripheral-heavy systems. It can; the difference is that its I/O is distributed across the board and must be matched to the node placement.
Conclusion
The NVIDIA Jetson Orin Nano Super Developer Kit and Turing Pi 2.5 solve different problems around the same module, but the difference is not as simple as “development board versus headless cluster.”
The developer kit is the cleanest single-Jetson package. It puts NVIDIA’s direct CSI, Jetson header, USB, display, wireless, power, and reference-carrier workflow around one module with minimal setup.
Turing Pi 2.5 keeps the Jetson’s compute essentially unchanged while giving it a much larger system to live in. Alongside per-node NVMe, integrated networking, centralized power, BMC control, and three neighboring compute slots, the board can also provide HDMI, USB, GPIO/I²C, and Mini PCIe expansion depending on slot placement. That makes it useful not only for homelabs and AI servers, but also for managed edge devices, vision systems, robotics, kiosks, and other builds where the Jetson is one part of a broader machine.
The measurements back that up. With the same module, NVMe, software, and power mode, LLM decode measured 25.276 tok/s on the developer kit and 25.215 tok/s on Turing Pi (a 0.24% difference inside run-to-run variance), neither carrier lost throughput over 30 minutes of sustained load, and median cold model loading, sequential reads, 4K random I/O, and GbE throughput all landed within a few percent. Two results differed: the module ran 8.7 °C cooler on the open-air developer kit, which likely reflects the different airflow and enclosure conditions but was not isolated, and 60-second sequential NVMe writes averaged 56% lower on the developer kit. A follow-up showed the developer kit starts at the same ~2,700 MiB/s and then drops, so the behavior is consistent with drive state and timing rather than a fixed carrier limit, and we do not attribute the gap to the carrier, although we did not capture a time series on Turing Pi to confirm how its run differed.
For a project built around the developer kit’s two direct CSI connectors or its exact all-on-one I/O layout, NVIDIA’s carrier remains the easier choice. For a system that needs the same Jetson performance plus storage, management, networking, expansion, local I/O, and room for other compute nodes, Turing Pi 2.5 gives the module a much more capable environment around it.
Related articles
- NVIDIA Jetson Orin Nano Super on Turing Pi 2.5: Complete Setup Guide
- NVIDIA Jetson on Turing Pi 2.5: Supported Modules and What You Can Build
- Local AI on Turing Pi with NVIDIA Jetson: When Edge AI Makes Sense
- NVIDIA Jetson Software Stack on Turing Pi 2.5: JetPack, CUDA, TensorRT & Containers Explained
- Preparing NVIDIA Jetson as an AI Node on Turing Pi 2.5
- Run Ollama on Jetson Orin Nano: Build a Local AI Server on Turing Pi 2.5
- What LLMs Actually Fit on a Jetson Orin Nano 8GB? Models, Context, and Runtimes Tested
FAQ
Does Turing Pi 2.5 make Jetson Orin Nano faster?
Not inherently. The Orin Nano’s CPU, GPU, Tensor Cores, and LPDDR5 memory live on the module. In our controlled test with the same module, NVMe, software, and 25W power mode, decode measured 25.276 tok/s on the developer kit and 25.215 tok/s on Turing Pi, a 0.24% difference inside run-to-run variance. Carrier choice can still affect sustained performance indirectly through power delivery, cooling, I/O, and system configuration, but neither carrier lost throughput over a 30-minute run.
Can I move the Orin Nano module from the developer kit to Turing Pi 2.5?
Yes. That is the module used throughout this series. Our 8GB P3767-0005 development module was removed from NVIDIA’s developer kit and installed in Turing Pi 2.5, where it was flashed, booted from NVMe, and used for CUDA and LLM workloads. We later moved it back to the developer kit for this comparison.
Can I move the same NVMe between both carriers?
Physically, yes, and in our test the installation booted unchanged on the developer kit with the same IP address, power mode, and NVMe PCIe address, and no reflash. That worked because our Turing Pi install was flashed with the dev-kit-super configuration. Carrier configuration and Jetson boot firmware can matter, so if your install used a different configuration, be prepared to reflash, and reproduce the same Jetson Linux and JetPack environment before measuring performance.
Can the Turing Pi BMC flash a Jetson by itself?
Not in the same sense as flashing a generic raw disk image. Turing Pi can place and route the Jetson through the recovery USB path so an external host can detect NVIDIA APX. NVIDIA’s Jetson flashing tools still run on that host and perform the actual Jetson Linux flash.
Does Turing Pi 2.5 provide remote power control for the Jetson?
Yes, and it was re-verified on the comparison system rather than assumed. On BMC firmware 2024.05.1 (daemon 2.3.4) a per-node power-off stopped the Jetson in under 5 seconds, a power-on brought it back to SSH in about 98 seconds with the same NVMe installation, and a per-node reset rebooted it in about 85 seconds. The caveats are the node-index conventions (power is 1-based, uart/reset/usb are 0-based, and a wrong index silently hits a different node) and the fact that the BMC reports power state, not wattage.
Which platform is better for cameras and robotics prototyping?
NVIDIA’s developer kit is the simpler choice when the project specifically depends on its two direct MIPI CSI connectors, the developer kit’s Jetson header layout, or several peripherals attached to the same carrier. Turing Pi 2.5 can still support substantial robotics and edge I/O: Node 1 provides HDMI, USB, and a 40-pin GPIO path, the board exposes additional I²C/GPIO and Mini PCIe expansion, and other slots add their own USB and storage options. Turing Pi documents a Jetson-specific HDMI switch configuration, although we did not test that display path because our Jetson was in Node 2. It becomes especially useful when perception or inference is only one part of the robot and neighboring nodes can handle control, logging, networking, storage, or other services. The deciding factor is the I/O topology you need, not whether the project is “robotics.”
Which platform is cheaper for one Jetson?
The developer kit is the simpler one-node purchase. NVIDIA currently lists the Orin Nano Super Developer Kit at $399, while Turing Pi 2.5 alone is $279 before adding a Jetson module and the rest of the system. Turing Pi’s value is the shared multi-node platform rather than a lower-cost single carrier.
References
- NVIDIA Jetson FAQ: https://developer.nvidia.com/embedded/faq
- NVIDIA JetPack 7.2 downloads and notes: https://developer.nvidia.com/embedded/jetpack/downloads
- NVIDIA Jetson Linux 39.2.1 Developer Guide: https://docs.nvidia.com/jetson/archives/r39.2.1/DeveloperGuide/
- NVIDIA Jetson Orin Nano Developer Kit hardware layout: https://docs.nvidia.com/jetson/orin-nano-devkit/user-guide/hardware_layout.html
- NVIDIA Jetson Orin Nano Developer Kit quick start: https://docs.nvidia.com/jetson/orin-nano-devkit/user-guide/quick_start.html
- NVIDIA Jetson developer kits: https://developer.nvidia.com/embedded/jetson-developer-kits
- Turing Pi 2.5 product page: https://turingpi.com/product/turing-pi-2-5/
- Turing Pi specs and I/O: https://docs.turingpi.com/docs/turing-pi2-specs-and-io-ports
- Turing Pi BMC API: https://docs.turingpi.com/docs/turing-pi2-bmc-api
- Turing Pi
tpiCLI usage: https://docs.turingpi.com/docs/tpi-usage - Turing Pi 2.5 changelog: https://docs.turingpi.com/changelog/turing-pi2-v25-list-of-improvements