Key Specifications

Vendorgoogle
Version1.5-pro
Release Date2024-02-15
Context Window2e+06 tokens
Input Modalitiestext, image, audio, video
Output Modalitiestext
LicenseProprietary
Documentationhttps://ai.google.dev/gemini-api/docs

Benchmark Performance

BenchmarkScoreUnitEvaluated AtNotesSource
MMLU85.9%2024-02-155-shotview
HUMANEVAL71.9pass@12024-02-15view
GSM8K91.7%2024-02-150-shot CoTview
MATH58.5%2024-02-150-shot CoTview
BBH84%2024-02-153-shot CoTview

Pricing

TierPriceCurrency
Input$1.25 / MtokUSD
Output$5 / MtokUSD
Cache Read$0.3125 / MtokUSD
Cache Write$1.25 / MtokUSD

Source: https://ai.google.dev/pricing · as of 2024-08-01

Compliance

  • Data Residency: US
  • SOC2: ✓
  • HIPAA: ✗
  • GDPR: ✓
  • ISO 27001: ✓

Gemini 1.5 Pro

Model Overview

Gemini 1.5 Pro is Google DeepMind’s flagship multimodal model, released on February 15, 2024. Its standout feature is the industry-leading 2 million token context window, enabling whole-codebase reasoning, multi-hour video understanding, and million-word document analysis in a single call. The model accepts text, image, audio, and video inputs natively, making it the only flagship that processes video as a first-class modality. Built on a mixture-of-experts architecture, Gemini 1.5 Pro delivers strong reasoning performance at roughly half the price of GPT-4o. It powers the Gemini Advanced consumer product, Google Cloud Vertex AI, and Google Workspace AI features. The long-context capability is particularly valuable for enterprises dealing with large legal contracts, financial filings, or research corpora.

Key Specifications

AttributeValue
VendorGoogle DeepMind
Version1.5-pro
Release Date2024-02-15
Context Window2,000,000 tokens
Input Modalitiestext, image, audio, video
Output Modalitiestext
LicenseProprietary
Documentationhttps://ai.google.dev/gemini-api/docs

Benchmark Performance

BenchmarkScoreUnitNotes
MMLU85.9%5-shot
HumanEval71.9pass@1
GSM8K91.7%0-shot CoT
MATH58.5%0-shot CoT
BBH84.0%3-shot CoT

Pricing

TierPrice (per 1M tokens)
Input$1.25
Output$5.00
Cache Read$0.3125
Cache Write$1.25

Pricing source: https://ai.google.dev/pricing (as of 2024-08-01).

Strengths

  • Industry-leading 2M token context window for ultra-long inputs.
  • Native video modality support, unique among flagship models.
  • Competitive pricing — roughly half of GPT-4o.
  • Strong multimodal grounding across text, image, audio, and video.

Weaknesses

  • HumanEval score (71.9) trails GPT-4o and Claude 3.5 Sonnet on pure coding.
  • Long-context recall can be inconsistent for facts buried deep in the input.
  • Proprietary model with no self-host option.

Use Cases

  • Whole-codebase refactoring and analysis.
  • Long-video summarization and content moderation.
  • Multi-document legal and financial research.
  • Cross-modal retrieval over mixed media archives.

References