Co-Founder & Lead Programmer of AcceleratedLogic AI
NVIDIA released Nemotron 3.5 Lightning on August 11, 2026. This page replaces an overbroad speed-focused review with a source-grounded summary of the model card and a transparent note about one saved code-generation sample. NVIDIA's reported benchmark scores are vendor evaluations; they are not results from an independent AcceleratedLogic benchmark.
Specifications and deployment context
NVIDIA's model card describes a 30-billion-parameter model with 3 billion active parameters, a hybrid Mamba-2/MoE/attention architecture, and up to a one-million-token context window. The card lists an NVFP4 checkpoint and the OpenMDW 1.1 license. The model is designed for NVIDIA GPU-accelerated systems; 3B active parameters does not mean the complete 30B checkpoint fits in 3B-parameter memory.
NVIDIA publishes several precision-specific checkpoints and deployment paths. Hardware needs depend on the selected checkpoint, serving configuration, context length, and concurrency. The NVIDIA NIM guide and model card are more useful for planning a local deployment than a generic claim that the model is easy to run on a personal computer.
Reading NVIDIA's benchmark results
The NVIDIA model card lists results for multiple evaluations, including MMLU Pro, GPQA Diamond, SWE-bench Verified, Terminal-Bench 2.1, and long-context tests. It distinguishes BF16 and NVFP4 checkpoints and links to evaluation recipes. These are NVIDIA-reported results using the stated checkpoints and harnesses; compare them with external results only after checking model precision, tools, prompt, harness, and scoring rules.
What one saved coding sample showed
An earlier site test asked the model to generate a self-contained Three.js globe. The stored output in that test was not executable as written. A few visible defects were missing arithmetic operators in expressions such as (i / particleCount) Math.PI 2, particlePositions[i 3], and orbitRadius Math.sin(polar). This observation applies to that saved response, not to all Nemotron coding tasks.
The original test also includes an SVG illustration that was judged visually weak against the requested crane prompt. That is a subjective result from one run, without a repeatable test harness or scored rubric. This rewrite therefore does not turn the sample into a general model rating.
How to test it for your workload
If you have compatible NVIDIA hardware, compare the BF16 and quantized checkpoints using tasks drawn from your own application. Save the model version, precision, GPU, serving parameters, prompt, generated output, run count, latency, and task-specific score. For code, parse and run the result and use tests; for text, score against a written rubric. Report hardware and harness with any benchmark so readers can understand what the result measures.