A 30B Open Agent Model Matched a 120B Peer. Your Route May Be Stale.

NVIDIA's Nemotron 3.5 Lightning scored 24 on Artificial Analysis's Intelligence Index, matching gpt-oss-120b at a quarter of the parameters. Remap the workhorse.

calender-image
August 14, 2026
clock-image
6 min read
A 30B Open Agent Model Matched a 120B Peer. Your Route May Be Stale.
Free weekly briefingThe Business AI Briefing for people who run the Business — 5 min, zero hype.
Get the briefing free →

Via Artificial Analysis: NVIDIA launches Nemotron 3.5 Lightning

A vendor-neutral scoreboard just moved the cheap agent slot

NVIDIA released Nemotron 3.5 Lightning on August 11, 2026. Artificial Analysis, an independent benchmarking shop, scored it 24 on its Intelligence Index — a nine-point jump over Nemotron 3 Nano (15). That puts Lightning in line with OpenAI's gpt-oss-120b (also 24) and only just behind NVIDIA's own Nemotron 3 Super (26), a model about four times its size.

The architecture is a hybrid Mamba-Transformer mixture of experts: 31.6 billion total parameters, 3.6 billion active, one million tokens of context, text-only reasoning. License: OpenMDW-1.1, open for commercial use without material restrictions, according to Artificial Analysis. Weights ship in BF16 and NVFP4; the quantized NVFP4 variant also scored 24 on the Index.

For an owner-led firm, the headline is not “NVIDIA won.” It is whether last quarter's model route — the default in Cursor, the API ID in a Zap, the “use the big one for agents” rule of thumb — still matches the jobs you actually run. Scoreboards move. Defaults do not, unless someone owns them.

Why Lightning matters now for SME model maps

Artificial Analysis's useful frame is not raw IQ. It is time and cost per task. On a pre-release DeepInfra endpoint serving the final NVFP4 weights, they measured median output speeds of nearly 670 tokens per second and about 0.5 minutes per Intelligence Index task. Peers in the same write-up were slower: Qwen3.6 35B A3B ~3.5 minutes, gpt-oss-120b ~3.4, Gemma 4 31B ~5.8, Qwen3.6 27B ~7.3.

The biggest gains over Nano sit on agentic evals. GDPval-AA v2 Elo of 824 moved past gpt-oss-120b and Nemotron 3 Super. Terminal-Bench v2.1 scored 24% versus Nano's 7% — more than a 3× jump, nearly matching gpt-oss-120b. Combined with speed and a permissive license, Artificial Analysis calls Lightning an efficient workhorse for high-volume agentic deployments.

It is not the smartest small model. Qwen3.6 35B A3B (32) and Muse Glimmer high (35) still lead the size class on the Index. Proprietary models still sit higher on the time-efficiency frontier: Gemini 3.5 Flash-Lite at 37 Intelligence with similar time per task, GPT-5.6 Luna (max) at 52 in under two minutes. Lightning is a different point on the frontier: cheaper active parameters, very fast tokens, agentic scores that just jumped.

Serverless inference is already listed from DeepInfra, Fireworks, FriendliAI, CoreWeave, GMI Cloud, Nebius, and Crusoe. That is availability. It is not a continuity plan. Endpoints, quantization, and prices will move. Your named primary and fallback should be written down before a provider changes the SKU.

Blog Image

What smart firms do when a workhorse model shifts

  1. Separate the jobs. High-volume agent loops (classify, extract, draft, tool-call) are not the same as partner-level judgment. Lightning is a candidate for the first pile. It is not an argument to put every client memo on the cheapest fast model.
  2. Re-run your own tasks. Artificial Analysis's Index is nine evaluations. Your inbox, your CRM notes, and your codebase are not those nine. Time a week of real jobs at current prices before you change the default.
  3. Name a primary and a fallback. If Lightning is the new workhorse, what fails over when NVFP4 serving blips, the license terms change, or a better Qwen lands? Write the ID, the provider, and the trigger.
  4. Watch token economics, not just tokens per second. Fast output is wasted if the model is verbose, or if staff send frontier work to a 3.6B-active workhorse and then rewrite it. Measure cost per accepted task, not cost per million tokens in a vacuum.
  5. Do not confuse open weights with no ops. OpenMDW-1.1 is commercially usable per the write-up. Serving, evals, and data-handling rules are still work. If nobody owns that work, you will drift back to whatever the IDE suggested this morning.

This puts it in line with OpenAI's gpt-oss-120b (24) and only just behind Nemotron 3 Super (26), a model ~4x its size. — Artificial Analysis

How AgentsROI maps this to Model Selection

AgentsROI.ai is a managed AI services provider for owner-led SMEs. We do not sell NVIDIA, Qwen, or OpenAI. We match the model to the job — with a fallback so a discontinued, restricted, or repriced SKU does not break the week.

Lead with Model Selection & Continuity Planning. Lightning's +9 Intelligence Index jump and agentic Elo move are a prompt to refresh the map: which workflows get a fast open workhorse, which stay on a proprietary frontier model, and what you pay per accepted task either way. We sell the judgment. Config can land on whichever provider is actually up.

Managed AI Operations is the destination once a route exists: monitor quality, spend, and silent default changes when Cursor, OpenRouter, or a Zap swaps the model ID. Fractional AI Officer is the closer for firms at the 50–100-employee end that need someone to own the operating tempo rather than a one-off bake-off.

Vendor-neutral means we will tell you if Qwen or a closed Flash-class model still wins your jobs. A scoreboard is a starting point. Your holdout set is the decision.

Update the workhorse before the defaults do it for you

A 30B-class open model matching a 120B peer on an independent index, at roughly 670 tokens per second in one pre-release test, is a procurement signal. It is not a reason to rip out everything by Friday. It is a reason to write down which agent jobs still belong on last quarter's ID.

If your team is already running agents off an unnamed default, start with Model Selection & Continuity Planning. Put Lightning, its peers, and your frontier fallback on one page with prices and a re-eval date. Book a no-pressure assessment when you want that map owned, not hoped for.

This article summarizes publicly reported information and is for general informational purposes only. It does not constitute legal, tax, financial, investment, security, or compliance advice. AgentsROI.ai is not a law firm, accounting firm, or registered investment adviser. Facts, pricing, statistics, and product capabilities cited here reflect the sources listed at the time of writing and may change. Readers should verify current information independently and consult qualified professionals regarding obligations specific to their industry, jurisdiction, and circumstances—including applicable New York State and New York City requirements. AgentsROI.ai may have commercial relationships with vendors mentioned; where material, such relationships are disclosed. Nothing in this article is an endorsement of any specific AI product, model, or provider.