Key Specifications
| Vendor | alibaba |
|---|
| Version | 2.5-72b |
|---|
| Release Date | 2024-09-19 |
|---|
| Context Window | 131072 tokens |
|---|
| Input Modalities | text |
|---|
| Output Modalities | text |
|---|
| License | Qwen License |
|---|
| Documentation | https://qwenlm.github.io/ |
|---|
Benchmark Performance
| Benchmark | Score | Unit | Evaluated At | Notes | Source |
|---|
| MMLU | 86.1 | % | 2024-09-19 | 5-shot | view |
| HUMANEVAL | 86.6 | pass@1 | 2024-09-19 | — | view |
| GSM8K | 88.4 | % | 2024-09-19 | 0-shot CoT | view |
| MATH | 83.1 | % | 2024-09-19 | 0-shot CoT | view |
| BBH | 82.4 | % | 2024-09-19 | 3-shot CoT | view |
Compliance
- Data Residency: CN
- SOC2: ✗
- HIPAA: ✗
- GDPR: ✗
- ISO 27001: ✗
Qwen2.5 72B
Ikhtisar Model
阿里巴巴通义千问 Qwen2.5 72B 开源模型,131K 上下文窗口,在中文理解、代码生成与数学推理上表现突出,是当前最强的中文开源模型之一。
Spesifikasi Inti
| Vendor | Versi | Tanggal Rilis | Jendela Konteks | Modalitas Input | Modalitas Output | Lisensi |
|---|
| Alibaba | 2.5-72b | 2024-09-19 | 131K | text | text | Qwen License |
Kinerja Benchmark
| Benchmark | Skor | Satuan | Catatan |
|---|
| MMLU (Massive Multitask Language Understanding) | 86.1 | % | 5-shot |
| HumanEval | 86.6 | pass@1 | — |
| GSM8K (Grade School Math 8K) | 88.4 | % | 0-shot CoT |
| MATH | 83.1 | % | 0-shot CoT |
| BBH (BIG-Bench Hard) | 82.4 | % | 3-shot CoT |
Harga
| Input | Output | Baca Cache | Tulis Cache |
|---|
| — | — | — | — |
per juta token
Kelebihan
- MMLU score 86.1, strong knowledge reasoning.
- HumanEval 86.6, excellent code generation.
- GSM8K 88.4, robust math reasoning.
Kekurangan
Kasus Penggunaan
Referensi