Key Specifications

Vendorother
Versionphi-1-5
Release Date2023-09-14
Context Window2048 tokens
Input Modalitiestext
Output Modalitiestext
LicenseMIT
Documentationhttps://huggingface.co/models

Benchmark Performance

BenchmarkScoreUnitEvaluated AtNotesSource
MMLU58.4%2023-09-145-shotview
HUMANEVAL39.2pass@12023-09-14view
GSM8K34.2%2023-09-140-shot CoTview
MATH27.2%2023-09-140-shot CoTview
BBH49.5%2023-09-143-shot CoTview
GPQA16.9%2023-09-140-shotview
IFEVAL48%2023-09-14prompt_strictview
ARC75.9%2023-09-14challengeview
MUSR30.3%2023-09-140-shotview
WINOGRANDE74.1%2023-09-140-shotview

Pricing

TierPriceCurrency
Input$0.1 / MtokUSD
Output$0.1 / MtokUSD
Cache Read$0 / MtokUSD
Cache Write$0 / MtokUSD

Source: https://huggingface.co/models · as of 2023-09-14

Compliance

  • Data Residency: self-host
  • SOC2: ✗
  • HIPAA: ✗
  • GDPR: ✗
  • ISO 27001: ✗

Phi-1.5

Model Overview

Microsoft Phi-1.5 1.3B 模型, 2K 上下文, 改进推理能力, 在 1.3B 规模上接近 7B 模型水平。

Core Specifications

VendorVersionRelease DateContext WindowInput ModalitiesOutput ModalitiesLicense
Otherphi-1-52023-09-142KtexttextMIT

Benchmark Performance

BenchmarkScoreUnitNotes
MMLU (Massive Multitask Language Understanding)58.4%5-shot
HumanEval39.2pass@1
GSM8K (Grade School Math 8K)34.2%0-shot CoT
MATH27.2%0-shot CoT
BBH (BIG-Bench Hard)49.5%3-shot CoT
GPQA16.9%0-shot
IFEval48.0%prompt_strict
ARC75.9%challenge
MUSR30.3%0-shot
WinoGrande74.1%0-shot

Pricing

InputOutputCache ReadCache Write

per million tokens

Strengths

  • Reliable general-purpose model.

Weaknesses

  • MMLU 58.4, weak knowledge reasoning.
  • HumanEval 39.2, coding weak.
  • Proprietary, not self-hostable.
  • Context window 2K is limited.

Use Cases

  • General chat and Q&A

References