Key Specifications

SpecificationPhi-3 VisionGPT-4o
Vendorotheropenai
Versionphi-3-vision4o
Release Date2024-05-212024-05-13
Context Window4096 tokens128000 tokens
Input Modalitiestext, imagetext, image, audio
Output Modalitiestexttext, audio
LicenseMITProprietary
SOC2
HIPAA
GDPR
ISO 27001

Benchmark Results

BenchmarkPhi-3 VisionGPT-4oWinner
ARC77.4Phi-3 Vision
BBH40.183.1GPT-4o
GPQA17.4Phi-3 Vision
GSM8K3095.8GPT-4o
HUMANEVAL48.290.2GPT-4o
IFEVAL56.4Phi-3 Vision
MATH24.376.6GPT-4o
MMLU56.488.7GPT-4o
MUSR32.1Phi-3 Vision
WINOGRANDE74.7Phi-3 Vision

Pricing Comparison

Tier (per Mtok)Phi-3 VisionGPT-4o
Input$0.2$2.5
Output$0.2$10
Cache Read$0$1.25
Cache Write$0$2.5

Phi-3 Vision vs GPT-4o

Ikhtisar Model

Phi-3 Vision and GPT-4o are both notable options in the AI model market. This page compares their benchmarks, pricing, and compliance.

Spesifikasi Utama

VendorTanggal RilisJendela KonteksLisensi
Other / Openai2024-05-21 / 2024-05-134K / 128KMIT / Proprietary

Kinerja Benchmark

BenchmarkPhi-3 VisionGPT-4oPemenang
ARC77.4A
BBH (BIG-Bench Hard)40.183.1B
GPQA17.4A
GSM8K (Grade School Math 8K)30.095.8B
HumanEval48.290.2B
IFEval56.4A
MATH24.376.6B
MMLU (Massive Multitask Language Understanding)56.488.7B
MUSR32.1A
WinoGrande74.7A

Perbandingan Harga

InputOutputBaca CacheTulis Cache
— / —— / —— / —— / —

per juta token — A / B

Kelebihan & Kekurangan

Phi-3 Vision

  • ✅ 支持文本、图像、音频多模态输入。
  • ⚠️ MMLU 仅 56.4,知识推理偏弱。
  • ⚠️ HumanEval 48.2,代码能力较弱。
  • ⚠️ 闭源专有模型,不支持自托管。
  • ⚠️ 上下文窗口 4K 偏小。

GPT-4o

  • ✅ MMLU score 88.7, strong knowledge reasoning.
  • ✅ HumanEval 90.2, excellent code generation.
  • ✅ GSM8K 95.8, robust math reasoning.
  • ✅ 支持文本、图像、音频多模态输入。
  • ⚠️ 闭源专有模型,不支持自托管。

Pendapat Editor

Phi-3 Vision and GPT-4o each have their strengths. Choose based on workload (code, long context, vision), referencing the tables above.

FAQ

Which model is better for coding tasks?

Refer to the HumanEval benchmark table; the model with a higher score is better suited for coding tasks.

Which model is cheaper?

Refer to the pricing comparison table above; the model with lower input/output prices is more cost-effective.

Which has a longer context window?

Refer to the key specifications table; the model with a larger context window is better for long documents.

Referensi

Editor's Take

See Editor's Take section.