核心规格对比

规格GPT-4oGemini 1.5 Pro
厂商openaigoogle
版本4o1.5-pro
发布日期2024-05-132024-02-15
上下文窗口128000 tokens2e+06 tokens
输入模态text, image, audiotext, image, audio, video
输出模态text, audiotext
许可ProprietaryProprietary
SOC2
HIPAA
GDPR
ISO 27001

基准测试结果

基准GPT-4oGemini 1.5 Pro胜出
BBH83.184Gemini 1.5 Pro
GSM8K95.891.7GPT-4o
HUMANEVAL90.271.9GPT-4o
MATH76.658.5GPT-4o
MMLU88.785.9GPT-4o

定价对比

项目 (每百万Token)GPT-4oGemini 1.5 Pro
输入$2.5$1.25
输出$10$5
缓存读$1.25$0.3125
缓存写$2.5$1.25

GPT-4o vs Claude 3.5 Sonnet — Comparison

Both models are flagship LLMs released in mid-2024. GPT-4o has stronger multimodal and audio capabilities, while Claude 3.5 Sonnet is slightly better at coding tasks.

Key Specifications

SpecificationGPT-4oClaude 3.5 Sonnet
VendorOpenAIAnthropic
Release Date2024-05-132024-06-20
Context Window128,000 tokens200,000 tokens
Input Modalitiestext, image, audiotext, image
Output Modalitiestext, audiotext

Performance

BenchmarkGPT-4oClaude 3.5 SonnetWinner
MMLU (%)88.788.7Tie
HumanEval (pass@1)90.292.0Claude wins
GSM8K (%)95.896.4Claude wins
MATH (%)76.671.1GPT-4o wins
BBH (%)83.184.5Claude wins

Pricing

GPT-4o: $2.50/$10 per Mtok (input/output). Claude: $3.00/$15 per Mtok. GPT-4o is ~30% cheaper overall on base pricing, while Claude 3.5 Sonnet offers substantially cheaper cache reads ($0.30 vs $1.25).

Strengths and Weaknesses

GPT-4o Strengths:

  • Native multimodal including audio input/output (unique among the two)
  • Lower base pricing on input and output tokens
  • HIPAA compliance for healthcare workloads
  • Lower latency, suitable for real-time voice

GPT-4o Weaknesses:

  • Smaller context window (128K vs 200K)
  • Trails Claude on HumanEval and BBH

Claude 3.5 Sonnet Strengths:

  • Industry-leading HumanEval (92.0 pass@1)
  • Larger 200K context window
  • Cheaper prompt cache reads ($0.30 vs $1.25)
  • Strong tool-use and agentic reliability

Claude 3.5 Sonnet Weaknesses:

  • No audio modality support
  • Higher base input/output pricing
  • No HIPAA compliance

Verdict

Choose Claude for coding and academic Q&A; choose GPT-4o for multimodal and general conversation.

编辑点评

## GPT-4o vs Claude 3.5 Sonnet — Comparison Both models are flagship LLMs released in mid-2024. GPT-4o has stronger multimodal and audio capabilities, while Claude 3.5 Sonnet is slightly better at coding tasks. ### Key Specifications | Specification | GPT-4o | Claude 3.5 Sonnet | |---------------|--------|--------------------| | Vendor | OpenAI | Anthropic | | Release Date | 2024-05-13 | 2024-06-20 | | Context Window | 128,000 tokens | 200,000 tokens | | Input Modalities | text, image, audio | text, image | | Output Modalities | text, audio | text | ### Performance | Benchmark | GPT-4o | Claude 3.5 Sonnet | Winner | |-----------|--------|--------------------|--------| | MMLU (%) | 88.7 | 88.7 | Tie | | HumanEval (pass@1) | 90.2 | 92.0 | Claude wins | | GSM8K (%) | 95.8 | 96.4 | Claude wins | | MATH (%) | 76.6 | 71.1 | GPT-4o wins | | BBH (%) | 83.1 | 84.5 | Claude wins | ### Pricing GPT-4o: $2.50/$10 per Mtok (input/output). Claude: $3.00/$15 per Mtok. GPT-4o is ~30% cheaper overall on base pricing, while Claude 3.5 Sonnet offers substantially cheaper cache reads ($0.30 vs $1.25). ### Strengths and Weaknesses **GPT-4o Strengths:** - Native multimodal including audio input/output (unique among the two) - Lower base pricing on input and output tokens - HIPAA compliance for healthcare workloads - Lower latency, suitable for real-time voice **GPT-4o Weaknesses:** - Smaller context window (128K vs 200K) - Trails Claude on HumanEval and BBH **Claude 3.5 Sonnet Strengths:** - Industry-leading HumanEval (92.0 pass@1) - Larger 200K context window - Cheaper prompt cache reads ($0.30 vs $1.25) - Strong tool-use and agentic reliability **Claude 3.5 Sonnet Weaknesses:** - No audio modality support - Higher base input/output pricing - No HIPAA compliance ### Verdict Choose Claude for coding and academic Q&A; choose GPT-4o for multimodal and general conversation.

常见问题

Placeholder question 1?
Replace with generated FAQ from body.
Placeholder question 2?
Replace with generated FAQ from body.