Piotr's picture

Piotr

piotr-ai

·

AI & ML interests

None yet

Recent Activity

liked a model 2 days ago

Infinigence/Megrez-3B-Instruct

liked a model 4 days ago

CohereForAI/c4ai-command-r7b-12-2024

liked a model 4 days ago

NexaAIDev/OmniAudio-2.6B

View all activity

Organizations

None yet

piotr-ai's activity

upvoted a paper 9 days ago

Transformers Can Navigate Mazes With Multi-Step Prediction

Paper • 2412.05117 • Published 12 days ago • 5

upvoted a collection 13 days ago

Common Models

The first generation of models pretrained on Common Corpus. • 5 items • Updated 13 days ago • 27

upvoted 3 collections about 2 months ago

SmolLM2

State-of-the-art compact LLMs for on-device applications: 1.7B, 360M, 135M • 15 items • Updated 16 days ago • 193

LayerSkip

Models continually pretrained using LayerSkip - https://arxiv.org/abs/2404.16710 • 8 items • Updated 27 days ago • 45

Granite 3.0 Language Models

A series of language models trained by IBM licensed under Apache 2.0 license. We release both the base pretrained and instruct models. • 8 items • Updated about 8 hours ago • 95

upvoted a paper 3 months ago

The AdEMAMix Optimizer: Better, Faster, Older

Paper • 2409.03137 • Published Sep 5 • 5

upvoted 6 collections 3 months ago

Llama 3.2

This collection hosts the transformers and original repos of the Llama 3.2 and Llama Guard 3 • 15 items • Updated 12 days ago • 543

Molmo

Artifacts for open multimodal language models. • 5 items • Updated 21 days ago • 288

Moshi v0.1 Release

MLX, Candle & PyTorch model checkpoints released as part of the Moshi release from Kyutai. Run inference via: https://github.com/kyutai-labs/moshi • 13 items • Updated Sep 18 • 224

Qwen2.5

Qwen2.5 language models, including pretrained and instruction-tuned models of 7 sizes, including 0.5B, 1.5B, 3B, 7B, 14B, 32B, and 72B. • 45 items • Updated 21 days ago • 426

DataGemma Release

A series of pioneering open models that help ground LLMs in real-world data through Data Commons. • 2 items • Updated 5 days ago • 81

Power-LM

Dense & MoE LLMs trained with power learning rate scheduler. • 4 items • Updated Oct 17 • 15

upvoted a paper 4 months ago

VideoLLaMB: Long-context Video Understanding with Recurrent Memory Bridges

Paper • 2409.01071 • Published Sep 2 • 27

upvoted 2 collections 4 months ago

Yi-Coder

4 items • Updated Sep 4 • 31

CogVLM2

This collection hosts the repos of the THUDM's CogVLM2 releases • 8 items • Updated 21 days ago • 19

upvoted a paper 4 months ago

CogVLM2: Visual Language Models for Image and Video Understanding

Paper • 2408.16500 • Published Aug 29 • 56

upvoted a collection 4 months ago

Qwen2-VL

Vision-language model series based on Qwen2 • 16 items • Updated 13 days ago • 180

upvoted 3 papers 4 months ago

LLaVA-MoD: Making LLaVA Tiny via MoE Knowledge Distillation

Paper • 2408.15881 • Published Aug 28 • 21

CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Paper • 2408.06072 • Published Aug 12 • 37

DeepSeek-Prover-V1.5: Harnessing Proof Assistant Feedback for Reinforcement Learning and Monte-Carlo Tree Search

Paper • 2408.08152 • Published Aug 15 • 52