Skip to content

Compare models

Put up to 4 models beside each other — token prices, context windows, capabilities and provider, from the same catalogue the model pages read.

  1. Nemotron 3 Nano Omni (Free)NVIDIARemove
  2. Gemini 3.8 FlashGoogleRemove
nemotron-3-nano-omni-30b-a3b-reasoning:free vs gemini-3.8-flash
AttributeNemotron 3 Nano Omni (Free)nemotron-3-nano-omni-30b-a3b-reasoning:freeGemini 3.8 Flashgemini-3.8-flash
Pricing
Input$0 / 1M$0.75 / 1M
Output$0 / 1M$3.75 / 1M
Cache Write$0 / 1M
Cache Read$0 / 1M$0.75 / 1M
Cache Write (5m)$0.75 / 1M
Cache Write (1h)$0.75 / 1M
Web Search$0 / 1M
Context
Max context256K1M
Max outputN/AN/A
Capabilities
VisionYesYes
Function CallingYesYes
JSON ModeYesYes
StreamingYesYes
Catalogue
ProviderNVIDIAGoogle
Categorychatchat
Charge typeFreePay As You Go
Released
Description
SummaryNVIDIA Nemotron 3 Nano Omni is an open 30B-A3B multimodal model designed as a perception and context sub-agent for enterprise agent systems. It supports text, image, video, and audio inputs with text output, enabling unified multimodal reasoning within a single inference loop. Built on a hybrid MoE Transformer–Mamba architecture with Conv3D video layers and Efficient Video Sampling (EVS), it delivers significantly improved efficiency for video reasoning—achieving ~2× higher throughput and 2.5× lower compute compared to separate pipelines. With up to 300K context length and extended thinking support, it is well suited for scalable, multimodal agent workflows.Gemini 3.8 Flash is Google's most intelligent Flash-class model, delivering significant improvements over Gemini 3.7 Flash across software engineering, agentic workflows, and complex multi-step reasoning. Designed to combine strong capability with Flash-tier efficiency, it is well suited for coding assistants, autonomous agents, and high-throughput production workflows that require responsive performance without sacrificing reasoning quality.