← All Models

NVIDIA: Llama 3.1 Nemotron 70B Instruct

🇺🇸 NVIDIA · Llama 3.1

Input Price per million tokens
Output Price per million tokens
Context Window 131K tokens Output limit: 16K
OpenRouter Route Price Please verify with official pricing pages
Use this model via OpenRouter →

Overview

NVIDIA: Llama 3.1 Nemotron 70B Instruct is a large language model API from NVIDIA, part of its Llama 3.1 model family. The 131K-token context window — around 197 pages of text — comfortably handles long documents, multi-file code, or extended conversations. On Artificial Analysis's Intelligence Index it scores 7 (F grade), a useful proxy for its general reasoning strength relative to the other models tracked here. All prices on this page reflect OpenRouter's routed rates and are re-synced automatically every day; confirm against the provider's official pricing before committing to production.

Dimension Unit Price (USD) Price (TWD) Effective From
No pricing data yet

Provider
NVIDIA
Model Family
Llama 3.1
Version String
nvidia/llama-3.1-nemotron-70b-instruct
Status
Active
Modality
Text
Context Window
131,072 tokens
Output Limit
16,384 tokens

Index Metrics

Cross-domain capability indexes evaluated by Artificial Analysis — Artificial Analysis

Intelligence Index 7 F Measured: 2026-08-29

Benchmark Scores

Data source: Artificial Analysis

AA-LCR 7.3% F Measured: 2026-08-29
GPQA Diamond 46.5% C Measured: 2026-08-29
IFBench 30.8% D Measured: 2026-08-29
Non-Hallucination 28.7% Measured: 2026-08-29
Omniscience Accuracy 17.8% Measured: 2026-08-29
SciCode 23.3% C Measured: 2026-08-29
Tau2 23.1% Measured: 2026-08-29
TerminalBench 4.5% Measured: 2026-08-29

Performance Metrics

Real-world benchmarks, updated every 72 hours by Artificial Analysis — Artificial Analysis

First Token Latency 13.3s Measured: 2026-08-29
Output Speed 36 t/s Measured: 2026-08-29
Response Time 27.0s Measured: 2026-08-29

Key Insights

Key data points from this page for quick reference and citation.

  • Context window: 131,072 tokens
  • Provider: NVIDIA
  • Model family: Llama 3.1
  • Modalities: Text
  • Data source: OpenRouter, updated daily