NFTCID

AI & ML interests

None yet

Recent Activity

liked a model 3 days ago

reacted to m-ric's post with 👀 13 days ago

𝗠𝗶𝗻𝗶𝗠𝗮𝘅'𝘀 𝗻𝗲𝘄 𝗠𝗼𝗘 𝗟𝗟𝗠 𝗿𝗲𝗮𝗰𝗵𝗲𝘀 𝗖𝗹𝗮𝘂𝗱𝗲-𝗦𝗼𝗻𝗻𝗲𝘁 𝗹𝗲𝘃𝗲𝗹 𝘄𝗶𝘁𝗵 𝟰𝗠 𝘁𝗼𝗸𝗲𝗻𝘀 𝗰𝗼𝗻𝘁𝗲𝘅𝘁 𝗹𝗲𝗻𝗴𝘁𝗵 💥 This work from Chinese startup @MiniMax-AI introduces a novel architecture that achieves state-of-the-art performance while handling context windows up to 4 million tokens - roughly 20x longer than current models. The key was combining lightning attention, mixture of experts (MoE), and a careful hybrid approach. 𝗞𝗲𝘆 𝗶𝗻𝘀𝗶𝗴𝗵𝘁𝘀: 🏗️ MoE with novel hybrid attention: ‣ Mixture of Experts with 456B total parameters (45.9B activated per token) ‣ Combines Lightning attention (linear complexity) for most layers and traditional softmax attention every 8 layers 🏆 Outperforms leading models across benchmarks while offering vastly longer context: ‣ Competitive with GPT-4/Claude-3.5-Sonnet on most tasks ‣ Can efficiently handle 4M token contexts (vs 256K for most other LLMs) 🔬 Technical innovations enable efficient scaling: ‣ Novel expert parallel and tensor parallel strategies cut communication overhead in half ‣ Improved linear attention sequence parallelism, multi-level padding and other optimizations achieve 75% GPU utilization (that's really high, generally utilization is around 50%) 🎯 Thorough training strategy: ‣ Careful data curation and quality control by using a smaller preliminary version of their LLM as a judge! Overall, not only is the model impressive, but the technical paper is also really interesting! 📝 It has lots of insights including a great comparison showing how a 2B MoE (24B total) far outperforms a 7B model for the same amount of FLOPs. Read it in full here 👉 https://huggingface.co/papers/2501.08313 Model here, allows commercial use <100M monthly users 👉 https://huggingface.co/MiniMaxAI/MiniMax-Text-01

liked a Space 15 days ago

akhaliq/anychat

View all activity

Organizations

None yet

NFTCID's activity

liked a model 3 days ago

deepseek-ai/DeepSeek-R1

Text Generation • Updated 5 days ago • 498k • 5.37k

liked a Space 15 days ago

Running on CPU Upgrade

1.65k

🏢

Anychat

liked 2 models 27 days ago

ibm-granite/granite-3.1-8b-instruct

Text Generation • Updated about 2 hours ago • 72.1k • 131

PowerInfer/SmallThinker-3B-Preview

Text Generation • Updated 15 days ago • 110k • 377

liked a dataset 27 days ago

agibot-world/AgiBotWorld-Alpha

Viewer • Updated 11 days ago • 19.7M • 22.6k • 165

liked a model 3 months ago

genmo/mochi-1-preview

Text-to-Video • Updated Dec 18, 2024 • 40.9k • 1.16k

liked a model 6 months ago

black-forest-labs/FLUX.1-schnell

Text-to-Image • Updated Aug 16, 2024 • 795k • 3.3k

liked a model about 1 year ago

ibm-research/re2g-reranker-trex

Text Classification • Updated May 16, 2023 • 1.64k • 7

liked a dataset about 1 year ago

Yelp/yelp_review_full

Viewer • Updated Jan 4, 2024 • 700k • 38.1k • 110