Key Specifications

SpecificationGPT-3.5 TurboGPT-4
Vendoropenaiopenai
Version3.5-turbo4
Release Date2023-03-012023-03-14
Context Window16385 tokens8192 tokens
Input Modalitiestexttext
Output Modalitiestexttext
LicenseProprietaryProprietary
SOC2
HIPAA
GDPR
ISO 27001

Benchmark Results

BenchmarkGPT-3.5 TurboGPT-4Winner
ARC86.894.1GPT-4
BBH48.280.9GPT-4
GPQA22.439.2GPT-4
GSM8K34.391.7GPT-4
HUMANEVAL51.377.6GPT-4
IFEVAL57.474.6GPT-4
MATH17.843.9GPT-4
MMLU61.185.8GPT-4
MUSR29.254.6GPT-4
WINOGRANDE76.780.3GPT-4

Pricing Comparison

Tier (per Mtok)GPT-3.5 TurboGPT-4
Input$0.5$30
Output$1.5$60
Cache Read$0$0
Cache Write$0$0

GPT-3.5 Turbo vs GPT-4

Model Overview

GPT-3.5 Turbo and GPT-4 are both notable options in the AI model market. This page compares their benchmarks, pricing, and compliance.

Key Specifications

VendorRelease DateContext WindowLicense
Openai / Openai2023-03-01 / 2023-03-1416K / 8KProprietary / Proprietary

Benchmark Performance

BenchmarkGPT-3.5 TurboGPT-4Winner
ARC86.894.1B
BBH (BIG-Bench Hard)48.280.9B
GPQA22.439.2B
GSM8K (Grade School Math 8K)34.391.7B
HumanEval51.377.6B
IFEval57.474.6B
MATH17.843.9B
MMLU (Massive Multitask Language Understanding)61.185.8B
MUSR29.254.6B
WinoGrande76.780.3B

Pricing Comparison

InputOutputCache ReadCache Write
— / —— / —— / —— / —

per million tokens — A / B

Strengths & Weaknesses

GPT-3.5 Turbo

  • ✅ Reliable general-purpose model.
  • ⚠️ Proprietary, not self-hostable.
  • ⚠️ Context window 16K is limited.

GPT-4

  • ✅ MMLU score 85.8, strong knowledge reasoning.
  • ✅ GSM8K 91.7, robust math reasoning.
  • ⚠️ Proprietary, not self-hostable.
  • ⚠️ Context window 8K is limited.

Editor’s Take

GPT-3.5 Turbo and GPT-4 each have their strengths. Choose based on workload (code, long context, vision), referencing the tables above.

FAQ

Which model is better for coding tasks?

Refer to the HumanEval benchmark table; the model with a higher score is better suited for coding tasks.

Which model is cheaper?

Refer to the pricing comparison table above; the model with lower input/output prices is more cost-effective.

Which has a longer context window?

Refer to the key specifications table; the model with a larger context window is better for long documents.

References

Editor's Take

See Editor's Take section.