Llama 3.3 70B Versatile on Groq
Production Groq text model for high-speed general chat and tool use.
- Input
- $0.59 / 1M tokens
- Output
- $0.79 / 1M tokens
- Cached read
- — / 1M tokens
- Cached write
- — / 1M tokens
- Batch discount
- —%
- Source
- Llama 3.3 70B Versatile on Groq pricing
- Verified
- Apr 2, 2026 (High)
Capabilities
- Modalities
- text→text
- Capabilities
- reasoningbatchSupportpromptCachingfunctionCallingstructuredOutputs
- Strengths
- Production-ready, Low-latency Groq hosting
- Tradeoffs
- Not a frontier reasoning model
Official Links
Benchmark Coverage
| Benchmark | Version | Score | Date | Source | Notes |
|---|
Release History
| Release | Alias | Lifecycle | Release Date | Deprecation | Shutdown | Summary |
|---|---|---|---|---|---|---|
| Llama 3.3 70B Versatile on Groq | groq-llama-3-3-70b-versatile | Active | Jan 6, 2025 | — | — | Current published model family snapshot. |
Host Coverage
| Host | Type | Context | Pricing Note | Differences |
|---|---|---|---|---|
| Groq API | first-party | 131.1K | Reference production Groq pricing. | Production model tier |
Migration Guidance
Default hosted Groq text tier when GPT-OSS depth is unnecessary.
Replacement models: groq-gpt-oss-120b
Change Events
| Date | Type | Title | Description | Source |
|---|---|---|---|---|
| Jan 6, 2025 | family_added | Llama 3.3 70B Versatile on Groq published | Initial public model family launch. | Llama 3.3 70B Versatile on Groq release notes |
Other models from Groq
GPT-OSS 120B on Groq, GPT-OSS 20B on Groq, Groq Compound, Llama 3.1 8B Instant on Groq, Llama 4 Scout on Groq, Qwen3 32B on Groq