所有对比
浏览 AI Benchmark Hub 上的所有模型对比。
AI Benchmark Hub 上的所有模型对比。
- ## GPT-4o vs Claude 3.5 Sonnet — Comparison
Both models are flagship LLMs released in mid-2024. GPT-4o has stronger multimodal and audio capabilities, while Claude 3.5 Sonnet is slightly better at coding tasks.
### Key Specifications | Specification | GPT-4o | Claude 3.5 Sonnet | |---------------|--------|--------------------| | Vendor | OpenAI | Anthropic | | Release Date | 2024-05-13 | 2024-06-20 | | Context Window | 128,000 tokens | 200,000 tokens | | Input Modalities | text, image, audio | text, image | | Output Modalities | text, audio | text |
### Performance | Benchmark | GPT-4o | Claude 3.5 Sonnet | Winner | |-----------|--------|--------------------|--------| | MMLU (%) | 88.7 | 88.7 | Tie | | HumanEval (pass@1) | 90.2 | 92.0 | Claude wins | | GSM8K (%) | 95.8 | 96.4 | Claude wins | | MATH (%) | 76.6 | 71.1 | GPT-4o wins | | BBH (%) | 83.1 | 84.5 | Claude wins |
### Pricing GPT-4o: $2.50/$10 per Mtok (input/output). Claude: $3.00/$15 per Mtok. GPT-4o is ~30% cheaper overall on base pricing, while Claude 3.5 Sonnet offers substantially cheaper cache reads ($0.30 vs $1.25).
### Strengths and Weaknesses
**GPT-4o Strengths:** - Native multimodal including audio input/output (unique among the two) - Lower base pricing on input and output tokens - HIPAA compliance for healthcare workloads - Lower latency, suitable for real-time voice
**GPT-4o Weaknesses:** - Smaller context window (128K vs 200K) - Trails Claude on HumanEval and BBH
**Claude 3.5 Sonnet Strengths:** - Industry-leading HumanEval (92.0 pass@1) - Larger 200K context window - Cheaper prompt cache reads ($0.30 vs $1.25) - Strong tool-use and agentic reliability
**Claude 3.5 Sonnet Weaknesses:** - No audio modality support - Higher base input/output pricing - No HIPAA compliance
### Verdict Choose Claude for coding and academic Q&A; choose GPT-4o for multimodal and general conversation.
## GPT-4o vs Claude 3.5 Sonnet — Comparison Both models are flagship LLMs released in mid-2024. GPT-4o has stronger multimodal and audio capabilities, while Claude 3.5 Sonnet is slightly better at coding tasks. ### Key Specifications | Specification | GPT-4o | Claude 3.5 Sonnet | |---------------|--------|--------------------| | Vendor | OpenAI | Anthropic | | Release Date | 2024-05-13 | 2024-06-20 | | Context Window | 128,000 tokens | 200,000 tokens | | Input Modalities | text, image, audio | text, image | | Output Modalities | text, audio | text | ### Performance | Benchmark | GPT-4o | Claude 3.5 Sonnet | Winner | |-----------|--------|--------------------|--------| | MMLU (%) | 88.7 | 88.7 | Tie | | HumanEval (pass@1) | 90.2 | 92.0 | Claude wins | | GSM8K (%) | 95.8 | 96.4 | Claude wins | | MATH (%) | 76.6 | 71.1 | GPT-4o wins | | BBH (%) | 83.1 | 84.5 | Claude wins | ### Pricing GPT-4o: $2.50/$10 per Mtok (input/output). Claude: $3.00/$15 per Mtok. GPT-4o is ~30% cheaper overall on base pricing, while Claude 3.5 Sonnet offers substantially cheaper cache reads ($0.30 vs $1.25). ### Strengths and Weaknesses **GPT-4o Strengths:** - Native multimodal including audio input/output (unique among the two) - Lower base pricing on input and output tokens - HIPAA compliance for healthcare workloads - Lower latency, suitable for real-time voice **GPT-4o Weaknesses:** - Smaller context window (128K vs 200K) - Trails Claude on HumanEval and BBH **Claude 3.5 Sonnet Strengths:** - Industry-leading HumanEval (92.0 pass@1) - Larger 200K context window - Cheaper prompt cache reads ($0.30 vs $1.25) - Strong tool-use and agentic reliability **Claude 3.5 Sonnet Weaknesses:** - No audio modality support - Higher base input/output pricing - No HIPAA compliance ### Verdict Choose Claude for coding and academic Q&A; choose GPT-4o for multimodal and general conversation.
- ## GPT-4o vs Claude 3.5 Sonnet — Comparison
Both models are flagship LLMs released in mid-2024. GPT-4o has stronger multimodal and audio capabilities, while Claude 3.5 Sonnet is slightly better at coding tasks.
### Key Specifications | Specification | GPT-4o | Claude 3.5 Sonnet | |---------------|--------|--------------------| | Vendor | OpenAI | Anthropic | | Release Date | 2024-05-13 | 2024-06-20 | | Context Window | 128,000 tokens | 200,000 tokens | | Input Modalities | text, image, audio | text, image | | Output Modalities | text, audio | text |
### Performance | Benchmark | GPT-4o | Claude 3.5 Sonnet | Winner | |-----------|--------|--------------------|--------| | MMLU (%) | 88.7 | 88.7 | Tie | | HumanEval (pass@1) | 90.2 | 92.0 | Claude wins | | GSM8K (%) | 95.8 | 96.4 | Claude wins | | MATH (%) | 76.6 | 71.1 | GPT-4o wins | | BBH (%) | 83.1 | 84.5 | Claude wins |
### Pricing GPT-4o: $2.50/$10 per Mtok (input/output). Claude: $3.00/$15 per Mtok. GPT-4o is ~30% cheaper overall on base pricing, while Claude 3.5 Sonnet offers substantially cheaper cache reads ($0.30 vs $1.25).
### Strengths and Weaknesses
**GPT-4o Strengths:** - Native multimodal including audio input/output (unique among the two) - Lower base pricing on input and output tokens - HIPAA compliance for healthcare workloads - Lower latency, suitable for real-time voice
**GPT-4o Weaknesses:** - Smaller context window (128K vs 200K) - Trails Claude on HumanEval and BBH
**Claude 3.5 Sonnet Strengths:** - Industry-leading HumanEval (92.0 pass@1) - Larger 200K context window - Cheaper prompt cache reads ($0.30 vs $1.25) - Strong tool-use and agentic reliability
**Claude 3.5 Sonnet Weaknesses:** - No audio modality support - Higher base input/output pricing - No HIPAA compliance
### Verdict Choose Claude for coding and academic Q&A; choose GPT-4o for multimodal and general conversation.
## GPT-4o vs Claude 3.5 Sonnet — Comparison Both models are flagship LLMs released in mid-2024. GPT-4o has stronger multimodal and audio capabilities, while Claude 3.5 Sonnet is slightly better at coding tasks. ### Key Specifications | Specification | GPT-4o | Claude 3.5 Sonnet | |---------------|--------|--------------------| | Vendor | OpenAI | Anthropic | | Release Date | 2024-05-13 | 2024-06-20 | | Context Window | 128,000 tokens | 200,000 tokens | | Input Modalities | text, image, audio | text, image | | Output Modalities | text, audio | text | ### Performance | Benchmark | GPT-4o | Claude 3.5 Sonnet | Winner | |-----------|--------|--------------------|--------| | MMLU (%) | 88.7 | 88.7 | Tie | | HumanEval (pass@1) | 90.2 | 92.0 | Claude wins | | GSM8K (%) | 95.8 | 96.4 | Claude wins | | MATH (%) | 76.6 | 71.1 | GPT-4o wins | | BBH (%) | 83.1 | 84.5 | Claude wins | ### Pricing GPT-4o: $2.50/$10 per Mtok (input/output). Claude: $3.00/$15 per Mtok. GPT-4o is ~30% cheaper overall on base pricing, while Claude 3.5 Sonnet offers substantially cheaper cache reads ($0.30 vs $1.25). ### Strengths and Weaknesses **GPT-4o Strengths:** - Native multimodal including audio input/output (unique among the two) - Lower base pricing on input and output tokens - HIPAA compliance for healthcare workloads - Lower latency, suitable for real-time voice **GPT-4o Weaknesses:** - Smaller context window (128K vs 200K) - Trails Claude on HumanEval and BBH **Claude 3.5 Sonnet Strengths:** - Industry-leading HumanEval (92.0 pass@1) - Larger 200K context window - Cheaper prompt cache reads ($0.30 vs $1.25) - Strong tool-use and agentic reliability **Claude 3.5 Sonnet Weaknesses:** - No audio modality support - Higher base input/output pricing - No HIPAA compliance ### Verdict Choose Claude for coding and academic Q&A; choose GPT-4o for multimodal and general conversation.
- ## GPT-4o vs Claude 3.5 Sonnet — Comparison
Both models are flagship LLMs released in mid-2024. GPT-4o has stronger multimodal and audio capabilities, while Claude 3.5 Sonnet is slightly better at coding tasks.
### Key Specifications | Specification | GPT-4o | Claude 3.5 Sonnet | |---------------|--------|--------------------| | Vendor | OpenAI | Anthropic | | Release Date | 2024-05-13 | 2024-06-20 | | Context Window | 128,000 tokens | 200,000 tokens | | Input Modalities | text, image, audio | text, image | | Output Modalities | text, audio | text |
### Performance | Benchmark | GPT-4o | Claude 3.5 Sonnet | Winner | |-----------|--------|--------------------|--------| | MMLU (%) | 88.7 | 88.7 | Tie | | HumanEval (pass@1) | 90.2 | 92.0 | Claude wins | | GSM8K (%) | 95.8 | 96.4 | Claude wins | | MATH (%) | 76.6 | 71.1 | GPT-4o wins | | BBH (%) | 83.1 | 84.5 | Claude wins |
### Pricing GPT-4o: $2.50/$10 per Mtok (input/output). Claude: $3.00/$15 per Mtok. GPT-4o is ~30% cheaper overall on base pricing, while Claude 3.5 Sonnet offers substantially cheaper cache reads ($0.30 vs $1.25).
### Strengths and Weaknesses
**GPT-4o Strengths:** - Native multimodal including audio input/output (unique among the two) - Lower base pricing on input and output tokens - HIPAA compliance for healthcare workloads - Lower latency, suitable for real-time voice
**GPT-4o Weaknesses:** - Smaller context window (128K vs 200K) - Trails Claude on HumanEval and BBH
**Claude 3.5 Sonnet Strengths:** - Industry-leading HumanEval (92.0 pass@1) - Larger 200K context window - Cheaper prompt cache reads ($0.30 vs $1.25) - Strong tool-use and agentic reliability
**Claude 3.5 Sonnet Weaknesses:** - No audio modality support - Higher base input/output pricing - No HIPAA compliance
### Verdict Choose Claude for coding and academic Q&A; choose GPT-4o for multimodal and general conversation.
## GPT-4o vs Claude 3.5 Sonnet — Comparison Both models are flagship LLMs released in mid-2024. GPT-4o has stronger multimodal and audio capabilities, while Claude 3.5 Sonnet is slightly better at coding tasks. ### Key Specifications | Specification | GPT-4o | Claude 3.5 Sonnet | |---------------|--------|--------------------| | Vendor | OpenAI | Anthropic | | Release Date | 2024-05-13 | 2024-06-20 | | Context Window | 128,000 tokens | 200,000 tokens | | Input Modalities | text, image, audio | text, image | | Output Modalities | text, audio | text | ### Performance | Benchmark | GPT-4o | Claude 3.5 Sonnet | Winner | |-----------|--------|--------------------|--------| | MMLU (%) | 88.7 | 88.7 | Tie | | HumanEval (pass@1) | 90.2 | 92.0 | Claude wins | | GSM8K (%) | 95.8 | 96.4 | Claude wins | | MATH (%) | 76.6 | 71.1 | GPT-4o wins | | BBH (%) | 83.1 | 84.5 | Claude wins | ### Pricing GPT-4o: $2.50/$10 per Mtok (input/output). Claude: $3.00/$15 per Mtok. GPT-4o is ~30% cheaper overall on base pricing, while Claude 3.5 Sonnet offers substantially cheaper cache reads ($0.30 vs $1.25). ### Strengths and Weaknesses **GPT-4o Strengths:** - Native multimodal including audio input/output (unique among the two) - Lower base pricing on input and output tokens - HIPAA compliance for healthcare workloads - Lower latency, suitable for real-time voice **GPT-4o Weaknesses:** - Smaller context window (128K vs 200K) - Trails Claude on HumanEval and BBH **Claude 3.5 Sonnet Strengths:** - Industry-leading HumanEval (92.0 pass@1) - Larger 200K context window - Cheaper prompt cache reads ($0.30 vs $1.25) - Strong tool-use and agentic reliability **Claude 3.5 Sonnet Weaknesses:** - No audio modality support - Higher base input/output pricing - No HIPAA compliance ### Verdict Choose Claude for coding and academic Q&A; choose GPT-4o for multimodal and general conversation.
- DeepSeek V3 vs Mistral Large 2:基准对比
See Editor's Take section.
- DeepSeek V3 vs Mixtral 8x22B:基准对比
See Editor's Take section.
- o1 vs DeepSeek V3:基准对比
See Editor's Take section.
- o1 vs Llama 3.1 405B:基准对比
See Editor's Take section.
- o1 vs o1 mini:基准对比
See Editor's Take section.
- Phi-4 vs Qwen2.5 14B:基准对比
See Editor's Take section.
- Gemini 2.0 Flash vs Gemini 2.0 Flash Thinking:基准对比
See Editor's Take section.
- Llama 3.3 70B vs Llama 3.1 70B:基准对比
See Editor's Take section.
- Claude 3.5 Haiku vs Gemini 1.5 Flash:基准对比
See Editor's Take section.
- Claude 3.5 Haiku vs Llama 3.1 8B:基准对比
See Editor's Take section.
- NVIDIA Llama 3.1 Nemotron 70B vs Llama 3.1 70B:基准对比
See Editor's Take section.
- Llama 3.2 11B Vision vs Gemma 2 9B:基准对比
See Editor's Take section.
- Llama 3.2 90B Vision vs GPT-4o:基准对比
See Editor's Take section.
- Qwen2.5 32B vs Mistral Small 3:基准对比
See Editor's Take section.
- Qwen2.5 32B vs Qwen2.5 14B:基准对比
See Editor's Take section.
- Qwen2.5 72B vs DeepSeek V3:基准对比
See Editor's Take section.
- Qwen2.5 72B vs Mistral Large 2:基准对比
See Editor's Take section.
- Qwen2.5 72B vs Mixtral 8x22B:基准对比
See Editor's Take section.
- o1 Preview vs DeepSeek V3:基准对比
See Editor's Take section.
- Grok-2 vs Claude 3 Opus:基准对比
See Editor's Take section.
- Grok-2 vs DeepSeek V3:基准对比
See Editor's Take section.
- Grok-2 vs Llama 3.1 405B:基准对比
See Editor's Take section.
- GLM-4 Plus vs DeepSeek V3:基准对比
See Editor's Take section.
- GLM-4 Plus vs Qwen2.5 72B:基准对比
See Editor's Take section.
- Hermes 3 Llama 3.1 405B vs Llama 3.1 405B:基准对比
See Editor's Take section.
- Jamba 1.5 Large vs Mistral Large 2:基准对比
See Editor's Take section.
- Mistral Large 2 vs Mistral Large:基准对比
See Editor's Take section.
- Llama 3.1 405B vs Command R+:基准对比
See Editor's Take section.
- Llama 3.1 405B vs DeepSeek V3:基准对比
See Editor's Take section.
- Llama 3.1 405B vs Mistral Large 2:基准对比
See Editor's Take section.
- Llama 3.1 405B vs Mixtral 8x22B:基准对比
See Editor's Take section.
- Llama 3.1 405B vs Qwen2.5 72B:基准对比
See Editor's Take section.
- Llama 3.1 70B vs DeepSeek V2:基准对比
See Editor's Take section.
- Llama 3.1 70B vs Mixtral 8x7B:基准对比
See Editor's Take section.
- Llama 3.1 70B vs Qwen2.5 72B:基准对比
See Editor's Take section.
- Llama 3.1 8B vs Mistral 7B v0.3:基准对比
See Editor's Take section.
- Llama 3.1 8B vs Qwen2.5 7B:基准对比
See Editor's Take section.
- GPT-4o mini vs Claude 3.5 Haiku:基准对比
See Editor's Take section.
- GPT-4o mini vs Gemini 1.5 Flash:基准对比
See Editor's Take section.
- GPT-4o mini vs Llama 3.1 8B:基准对比
See Editor's Take section.
- GPT-4o mini vs Mistral 7B v0.3:基准对比
See Editor's Take section.
- GPT-4o mini vs Qwen2.5 7B:基准对比
See Editor's Take section.
- Gemma 2 27B vs Gemma 2 9B:基准对比
See Editor's Take section.
- Gemma 2 9B vs Gemma 7B:基准对比
See Editor's Take section.
- DeepSeek Coder V2 vs Codestral:基准对比
See Editor's Take section.
- DeepSeek Coder V2 vs DeepSeek V3:基准对比
See Editor's Take section.
- Claude 3.5 Sonnet vs Claude 3 Opus:基准对比
See Editor's Take section.
- Claude 3.5 Sonnet vs DeepSeek V3:基准对比
See Editor's Take section.
- Claude 3.5 Sonnet vs Gemini 1.5 Pro:基准对比
See Editor's Take section.
- Claude 3.5 Sonnet vs Grok-2:基准对比
See Editor's Take section.
- Claude 3.5 Sonnet vs Qwen2.5 72B:基准对比
See Editor's Take section.
- Qwen2 72B vs Qwen2.5 72B:基准对比
See Editor's Take section.
- GLM-4 9B Chat vs Qwen2.5 7B:基准对比
See Editor's Take section.
- Phi-3 Medium vs Phi-3 Small:基准对比
See Editor's Take section.
- Phi-3 Vision vs GPT-4o:基准对比
See Editor's Take section.
- Gemini 1.5 Flash vs Llama 3.1 8B:基准对比
See Editor's Take section.
- GPT-4o vs Claude 3 Opus:基准对比
See Editor's Take section.
- GPT-4o vs Claude 3.5 Haiku:基准对比
See Editor's Take section.
- GPT-4o vs DeepSeek V3:基准对比
See Editor's Take section.
- GPT-4o vs Gemini 2.0 Flash:基准对比
See Editor's Take section.
- GPT-4o vs GPT-4 Turbo:基准对比
See Editor's Take section.
- GPT-4o vs GPT-4o (2024-08-06):基准对比
See Editor's Take section.
- GPT-4o vs GPT-4o mini:基准对比
See Editor's Take section.
- GPT-4o vs Grok-2:基准对比
See Editor's Take section.
- GPT-4o vs Llama 3.1 405B:基准对比
See Editor's Take section.
- GPT-4o vs Mistral Large 2:基准对比
See Editor's Take section.
- GPT-4o vs o1 Preview:基准对比
See Editor's Take section.
- GPT-4o vs o1:基准对比
See Editor's Take section.
- GPT-4o vs Qwen2.5 72B:基准对比
See Editor's Take section.
- Yi Large vs Qwen2.5 72B:基准对比
See Editor's Take section.
- Yi Large vs Yi 34B:基准对比
See Editor's Take section.
- DeepSeek V2 vs DeepSeek V3:基准对比
See Editor's Take section.
- Llama 3 70B vs Llama 3 8B:基准对比
See Editor's Take section.
- Mixtral 8x22B vs Mistral Large 2:基准对比
See Editor's Take section.
- DBRX Instruct vs Mixtral 8x22B:基准对比
See Editor's Take section.
- Command R vs Command R+:基准对比
See Editor's Take section.
- Claude 3 Haiku vs Claude 3 Sonnet:基准对比
See Editor's Take section.
- Claude 3 Opus vs Claude 3 Sonnet:基准对比
See Editor's Take section.
- StarCoder2 15B vs Code Llama 34B:基准对比
See Editor's Take section.
- Mistral Large vs Mistral Medium:基准对比
See Editor's Take section.
- Qwen1.5 72B vs Qwen2 72B:基准对比
See Editor's Take section.
- Gemini 1.5 Pro vs Claude 3 Opus:基准对比
See Editor's Take section.
- Gemini 1.5 Pro vs DeepSeek V3:基准对比
See Editor's Take section.
- Gemini 1.5 Pro vs Gemini 1.5 Flash:基准对比
See Editor's Take section.
- Gemini 1.5 Pro vs Gemini 2.0 Flash:基准对比
See Editor's Take section.
- Gemini 1.5 Pro vs Llama 3.1 405B:基准对比
See Editor's Take section.
- Gemini 1.5 Pro vs Qwen2.5 72B:基准对比
See Editor's Take section.
- Code Llama 70B vs Codestral:基准对比
See Editor's Take section.
- DeepSeek Coder 33B vs Code Llama 34B:基准对比
See Editor's Take section.
- Mixtral 8x7B vs Mixtral 8x22B:基准对比
See Editor's Take section.
- Gemini 1.0 Pro vs Gemini 1.0 Ultra:基准对比
See Editor's Take section.
- Claude 2.1 vs Claude 2:基准对比
See Editor's Take section.
- Yi 34B vs Yi 6B:基准对比
See Editor's Take section.
- Falcon 180B vs Llama 2 70B:基准对比
See Editor's Take section.
- Llama 2 70B vs Llama 3 70B:基准对比
See Editor's Take section.
- GPT-4 vs GPT-4 Turbo:基准对比
See Editor's Take section.
- GPT-3.5 Turbo vs GPT-4:基准对比
See Editor's Take section.