Hugging Face
Models
Datasets
Spaces
Posts
Docs
Enterprise
Pricing
Log In
Sign Up
SultanR
/
SmolTulu-1.7b-Reinforced
like
5
Text Generation
Transformers
Safetensors
allenai/RLVR-GSM-MATH-IF-Mixed-Constraints
English
llama
Tulu3
Smollm
SLMs
Small
Huggingface
Allenai
SFT
DPO
GGUF
RLVR
RL
conversational
text-generation-inference
Inference Endpoints
arxiv:
2411.15124
arxiv:
2412.08347
License:
apache-2.0
Model card
Files
Files and versions
Community
Train
Deploy
Use this model
530b6c0
SmolTulu-1.7b-Reinforced
/
smoltulubanner.png
SultanR
Upload smoltulubanner.png
7247435
verified
4 days ago
download
Copy download link
history
Safe
185 kB