Maharshi Gor's picture

9 3

Maharshi Gor

mgor

·

https://mgor.info

AI & ML interests

NLP, LLM Evaluations, Representation Learning, Question Answering

Recent Activity

upvoted a collection 14 days ago

Open LLM Leaderboard best models ❤️‍🔥

liked a Space 15 days ago

MAYA-AI/all-leaderboard

reacted to mayafree's post with 🔥 15 days ago

Leaderboard of Leaderboards — A Real-Time Meta-Ranking of AI Benchmarks https://huggingface.co/spaces/MAYA-AI/all-leaderboard Hundreds of AI leaderboards exist on HuggingFace. Knowing which ones the community actually trusts has never been easy — until now. Leaderboard of Leaderboards (LoL) ranks the leaderboards themselves, using live HuggingFace trending scores and cumulative likes as the signal. No editorial curation. No manual selection. Just what the global AI research community is actually visiting and endorsing, surfaced in real time. Sort by trending to see what is capturing attention right now, or by likes to see what has built lasting credibility over time. Nine domain filters let you zero in on what matters most to your work, and every entry shows both its rank within this collection and its real-time global rank across all HuggingFace Spaces. The collection spans well-established standards like Open LLM Leaderboard, Chatbot Arena, MTEB, and BigCodeBench alongside frameworks worth watching. FINAL Bench targets AGI-level evaluation across 100 tasks in 15 domains and recently reached the global top 5 in HuggingFace dataset rankings. Smol AI WorldCup runs tournament-format competitions for sub-8B models scored via FINAL Bench criteria. ALL Bench aggregates results across frameworks into a unified ranking that resists the overfitting risks of any single standard. The deeper purpose is not convenience. It is transparency. How we measure AI matters as much as the AI we measure.

View all activity

Organizations

upvoted a collection 14 days ago

Open LLM Leaderboard best models ❤️‍🔥

A daily uploaded list of models with best evaluations on the LLM leaderboard: • 50 items • Updated 15 days ago • 678

upvoted 2 collections 10 months ago

🧠 Reasoning model 2025

19 items • Updated Jan 4 • 6

🔥LLM 2025

6 items • Updated Jan 4 • 2

upvoted 2 papers about 1 year ago

R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model

Paper • 2503.05132 • Published Mar 7, 2025 • 57

R2-T2: Re-Routing in Test-Time for Multimodal Mixture-of-Experts

Paper • 2502.20395 • Published Feb 27, 2025 • 45

upvoted 2 papers over 1 year ago

BenTo: Benchmark Task Reduction with In-Context Transferability

Paper • 2410.13804 • Published Oct 17, 2024 • 20

Do great minds think alike? Investigating Human-AI Complementarity in Question Answering with CAIMIRA

Paper • 2410.06524 • Published Oct 9, 2024 • 4

upvoted a collection over 1 year ago

CAIMIRA Paper & Data

Question Answering Datasets on Quizbowl questions and their progressive clues from various competitions. • 3 items • Updated 26 days ago • 1

upvoted a paper almost 2 years ago

Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models

Paper • 2404.18796 • Published Apr 29, 2024 • 71