Dolphin 2.9.1 Yi 1.5 34B

Name: Dolphin 2.9.1 Yi 1.5 34B
Author: dphn

Downloads

Hugging Face

4.7M

HF Likes

Hugging Face

Context

GPT-3 class

License

—

Updated

11/3/2025

dphn

This model is based on the Yi 1.5 34B architecture and is licensed under Apache 2.0. It was generated from a trainer and utilizes various datasets including cognitivecomputations/Dolphin-2.9, teknium/OpenHermes-2.5, m-a-p/CodeFeedback-Filtered-Instruction, cognitivecomputations/dolphin-coder, cognitivecomputations/samantha-data, microsoft/orca-math-word-problems-200k, Locutusque/function-calling-chatml, and internlm/Agent-FLAN.

Language Model

OTHER

Try on Hugging Face

Add to Compare

Quick Info

Released

5/18/2024

Framework

OTHER

Resources

Training Data Analysis

🟡 Average (5.2/10)

Researched training datasets used by Dolphin 2.9.1 Yi 1.5 34B with quality assessment

Specialized For

code

general

science

multilingual

Training Datasets (3)

the pile

🟢 8/10

code

general

science

multilingual

Key Strengths

•Deliberate Diversity: Explicitly curated to include diverse content types (academia, code, Q&A, book...
•Documented Quality: Each component dataset is thoroughly documented with rationale for inclusion, en...
•Epoch Weighting: Component datasets receive different training epochs based on perceived quality, al...

common crawl

🔴 2.5/10

general

science

Key Strengths

•Scale and Accessibility: At 9.5+ petabytes, Common Crawl provides unprecedented scale for training d...
•Diversity: The dataset captures billions of web pages across multiple domains and content types, ena...
•Comprehensive Coverage: Despite limitations, Common Crawl attempts to represent the broader web acro...

Considerations

•Biased Coverage: The crawling process prioritizes frequently linked domains, making content from dig...
•Large-Scale Problematic Content: Contains significant amounts of hate speech, pornography, violent c...

wikipedia

🟡 5/10

science

multilingual

Key Strengths

•High-Quality Content: Wikipedia articles are subject to community review, fact-checking, and citatio...
•Multilingual Coverage: Available in 300+ languages, enabling training of models that understand and ...
•Structured Knowledge: Articles follow consistent formatting with clear sections, allowing models to ...

Considerations

•Language Inequality: Low-resource language editions have significantly lower quality, fewer articles...
•Biased Coverage: Reflects biases in contributor demographics; topics related to Western culture and ...

Explore our comprehensive training dataset analysis

View All Datasets