Key Specifications

Vendorxai
Version2
Release Date2024-08-13
Context Window131072 tokens
Input Modalitiestext, image
Output Modalitiestext
LicenseProprietary
Documentationhttps://docs.x.ai/

Benchmark Performance

BenchmarkScoreUnitEvaluated AtNotesSource
MMLU87.5%2024-08-135-shotview
HUMANEVAL88.4pass@12024-08-13view
GSM8K93.2%2024-08-130-shot CoTview
MATH76.8%2024-08-130-shot CoTview
BBH84%2024-08-133-shot CoTview

Pricing

TierPriceCurrency
Input$2 / MtokUSD
Output$10 / MtokUSD
Cache Read$0 / MtokUSD
Cache Write$0 / MtokUSD

Source: https://x.ai/api · as of 2024-08-13

Compliance

  • Data Residency: US
  • SOC2: ✗
  • HIPAA: ✗
  • GDPR: ✗
  • ISO 27001: ✗

Grok-2

Model Overview

Grok-2 is xAI’s flagship model, released on August 13, 2024. The model is notable for its tight integration with the X (formerly Twitter) platform, providing real-time awareness of current events and trending topics — a capability no other flagship model matches. Grok-2 accepts both text and image inputs, supports a 131K context window, and is positioned as a competitive alternative to GPT-4o and Claude 3.5 Sonnet on academic benchmarks. On MMLU, Grok-2 scores 87.5, within striking distance of GPT-4o (88.7) and Claude 3.5 Sonnet (88.7). The model is available via the X platform’s Grok chatbot (for Premium subscribers) and the xAI API. Grok-2 is also notable for its relatively permissive content policy compared to other flagship models, which xAI positions as a feature for users seeking fewer refusals.

Key Specifications

AttributeValue
VendorxAI
Version2
Release Date2024-08-13
Context Window131,072 tokens
Input Modalitiestext, image
Output Modalitiestext
LicenseProprietary
Documentationhttps://docs.x.ai/

Benchmark Performance

BenchmarkScoreUnitNotes
MMLU87.5%5-shot
HumanEval88.4pass@1
GSM8K93.2%0-shot CoT
MATH76.8%0-shot CoT
BBH84.0%3-shot CoT

Pricing

TierPrice (per 1M tokens)
Input$2.00
Output$10.00
Cache Read$0.00
Cache Write$0.00

Pricing source: https://x.ai/api (as of 2024-08-13).

Strengths

  • Real-time awareness via X platform integration, unique among flagships.
  • Strong benchmark performance, competitive with GPT-4o on most metrics.
  • Multimodal input (text + image) support.
  • More permissive content policy for users seeking fewer refusals.

Weaknesses

  • Proprietary model with no self-host option and limited enterprise compliance certifications.
  • No audio modality, unlike GPT-4o.
  • Tightly coupled to the X ecosystem, which may be a concern for some enterprise buyers.
  • Limited track record vs established vendors like OpenAI and Anthropic.

Use Cases

  • Real-time news and event-aware chatbots.
  • Social media content generation and trend analysis.
  • Customer-facing applications where current events matter.
  • Image-based question answering in conjunction with text.

References