Key Specifications

Vendoranthropic
Version3-opus
Release Date2024-03-04
Context Window200000 tokens
Input Modalitiestext, image
Output Modalitiestext
LicenseProprietary
Documentationhttps://docs.anthropic.com/claude/docs

Benchmark Performance

BenchmarkScoreUnitEvaluated AtNotesSource
MMLU81.2%2024-03-045-shotview
HUMANEVAL83.3pass@12024-03-04view
GSM8K91.8%2024-03-040-shot CoTview
MATH42.7%2024-03-040-shot CoTview
BBH81.6%2024-03-043-shot CoTview
GPQA46%2024-03-040-shotview
IFEVAL79.2%2024-03-04prompt_strictview
ARC94.5%2024-03-04challengeview
MUSR53.3%2024-03-040-shotview
WINOGRANDE81.9%2024-03-040-shotview

Pricing

TierPriceCurrency
Input$15 / MtokUSD
Output$75 / MtokUSD
Cache Read$0 / MtokUSD
Cache Write$0 / MtokUSD

Source: https://www.anthropic.com/pricing · as of 2024-03-04

Compliance

  • Data Residency: US
  • SOC2: ✓
  • HIPAA: ✗
  • GDPR: ✓
  • ISO 27001: ✗

Claude 3 Opus

Model Overview

Anthropic Claude 3 Opus 旗舰模型, 200K 上下文, 超越 GPT-4 的推理与创意能力, 适合复杂分析任务。

Core Specifications

VendorVersionRelease DateContext WindowInput ModalitiesOutput ModalitiesLicense
Anthropic3-opus2024-03-04200Ktext, imagetextProprietary

Benchmark Performance

BenchmarkScoreUnitNotes
MMLU (Massive Multitask Language Understanding)81.2%5-shot
HumanEval83.3pass@1
GSM8K (Grade School Math 8K)91.8%0-shot CoT
MATH42.7%0-shot CoT
BBH (BIG-Bench Hard)81.6%3-shot CoT
GPQA46.0%0-shot
IFEval79.2%prompt_strict
ARC94.5%challenge
MUSR53.3%0-shot
WinoGrande81.9%0-shot

Pricing

InputOutputCache ReadCache Write

per million tokens

Strengths

  • MMLU score 81.2, strong knowledge reasoning.
  • HumanEval 83.3, excellent code generation.
  • GSM8K 91.8, robust math reasoning.
  • Input Modalities: text, image, audio.
  • Context window 200K.

Weaknesses

  • Proprietary, not self-hostable.

Use Cases

  • Code generation and debugging
  • Long document summarization
  • Vision and image understanding
  • Agent workflows and tool use

References