NVIDIA's Nemotron 3.5 Lightning scored 24 on Artificial Analysis's Intelligence Index, matching gpt-oss-120b at a quarter of the parameters. Remap the workhorse.

Via Artificial Analysis: NVIDIA launches Nemotron 3.5 Lightning
NVIDIA released Nemotron 3.5 Lightning on August 11, 2026. Artificial Analysis, an independent benchmarking shop, scored it 24 on its Intelligence Index — a nine-point jump over Nemotron 3 Nano (15). That puts Lightning in line with OpenAI's gpt-oss-120b (also 24) and only just behind NVIDIA's own Nemotron 3 Super (26), a model about four times its size.
The architecture is a hybrid Mamba-Transformer mixture of experts: 31.6 billion total parameters, 3.6 billion active, one million tokens of context, text-only reasoning. License: OpenMDW-1.1, open for commercial use without material restrictions, according to Artificial Analysis. Weights ship in BF16 and NVFP4; the quantized NVFP4 variant also scored 24 on the Index.
For an owner-led firm, the headline is not “NVIDIA won.” It is whether last quarter's model route — the default in Cursor, the API ID in a Zap, the “use the big one for agents” rule of thumb — still matches the jobs you actually run. Scoreboards move. Defaults do not, unless someone owns them.
Artificial Analysis's useful frame is not raw IQ. It is time and cost per task. On a pre-release DeepInfra endpoint serving the final NVFP4 weights, they measured median output speeds of nearly 670 tokens per second and about 0.5 minutes per Intelligence Index task. Peers in the same write-up were slower: Qwen3.6 35B A3B ~3.5 minutes, gpt-oss-120b ~3.4, Gemma 4 31B ~5.8, Qwen3.6 27B ~7.3.
The biggest gains over Nano sit on agentic evals. GDPval-AA v2 Elo of 824 moved past gpt-oss-120b and Nemotron 3 Super. Terminal-Bench v2.1 scored 24% versus Nano's 7% — more than a 3× jump, nearly matching gpt-oss-120b. Combined with speed and a permissive license, Artificial Analysis calls Lightning an efficient workhorse for high-volume agentic deployments.
It is not the smartest small model. Qwen3.6 35B A3B (32) and Muse Glimmer high (35) still lead the size class on the Index. Proprietary models still sit higher on the time-efficiency frontier: Gemini 3.5 Flash-Lite at 37 Intelligence with similar time per task, GPT-5.6 Luna (max) at 52 in under two minutes. Lightning is a different point on the frontier: cheaper active parameters, very fast tokens, agentic scores that just jumped.
Serverless inference is already listed from DeepInfra, Fireworks, FriendliAI, CoreWeave, GMI Cloud, Nebius, and Crusoe. That is availability. It is not a continuity plan. Endpoints, quantization, and prices will move. Your named primary and fallback should be written down before a provider changes the SKU.
This puts it in line with OpenAI's gpt-oss-120b (24) and only just behind Nemotron 3 Super (26), a model ~4x its size. — Artificial Analysis
AgentsROI.ai is a managed AI services provider for owner-led SMEs. We do not sell NVIDIA, Qwen, or OpenAI. We match the model to the job — with a fallback so a discontinued, restricted, or repriced SKU does not break the week.
Lead with Model Selection & Continuity Planning. Lightning's +9 Intelligence Index jump and agentic Elo move are a prompt to refresh the map: which workflows get a fast open workhorse, which stay on a proprietary frontier model, and what you pay per accepted task either way. We sell the judgment. Config can land on whichever provider is actually up.
Managed AI Operations is the destination once a route exists: monitor quality, spend, and silent default changes when Cursor, OpenRouter, or a Zap swaps the model ID. Fractional AI Officer is the closer for firms at the 50–100-employee end that need someone to own the operating tempo rather than a one-off bake-off.
Vendor-neutral means we will tell you if Qwen or a closed Flash-class model still wins your jobs. A scoreboard is a starting point. Your holdout set is the decision.
A 30B-class open model matching a 120B peer on an independent index, at roughly 670 tokens per second in one pre-release test, is a procurement signal. It is not a reason to rip out everything by Friday. It is a reason to write down which agent jobs still belong on last quarter's ID.
If your team is already running agents off an unnamed default, start with Model Selection & Continuity Planning. Put Lightning, its peers, and your frontier fallback on one page with prices and a re-eval date. Book a no-pressure assessment when you want that map owned, not hoped for.
This article summarizes publicly reported information and is for general informational purposes only. It does not constitute legal, tax, financial, investment, security, or compliance advice. AgentsROI.ai is not a law firm, accounting firm, or registered investment adviser. Facts, pricing, statistics, and product capabilities cited here reflect the sources listed at the time of writing and may change. Readers should verify current information independently and consult qualified professionals regarding obligations specific to their industry, jurisdiction, and circumstances—including applicable New York State and New York City requirements. AgentsROI.ai may have commercial relationships with vendors mentioned; where material, such relationships are disclosed. Nothing in this article is an endorsement of any specific AI product, model, or provider.