Kanana: Compute-efficient Bilingual Language Models Paper • 2502.18934 • Published 9 days ago • 59
Running 2.1k 2.1k The Ultra-Scale Playbook 🌌 The ultimate guide to training LLM on large GPU Clusters
Mamba: Linear-Time Sequence Modeling with Selective State Spaces Paper • 2312.00752 • Published Dec 1, 2023 • 142
Rephrasing the Web: A Recipe for Compute and Data-Efficient Language Modeling Paper • 2401.16380 • Published Jan 29, 2024 • 49