Hugging Face
Models
Datasets
Spaces
Posts
Docs
Enterprise
Pricing
Log In
Sign Up
Audreygyj
/
pythia-160m-online-dpo-HH-2
like
0
Transformers
Safetensors
XueyingJia/online_dpo_repo
Generated from Trainer
trl
online-dpo
Inference Endpoints
arxiv:
2402.04792
Model card
Files
Files and versions
Community
Train
Deploy
Use this model
main
pythia-160m-online-dpo-HH-2
Commit History
End of training
7df72bd
verified
Audreygyj
commited on
Dec 9, 2024
Model save
3ab3300
verified
Audreygyj
commited on
Dec 9, 2024
Training in progress, step 8040
40f010b
verified
Audreygyj
commited on
Dec 9, 2024
Training in progress, step 7236
7eef100
verified
Audreygyj
commited on
Dec 9, 2024
Training in progress, step 6432
0bec8a7
verified
Audreygyj
commited on
Dec 8, 2024
Training in progress, step 5628
9ed52ef
verified
Audreygyj
commited on
Dec 8, 2024
Training in progress, step 4824
1f663b5
verified
Audreygyj
commited on
Dec 8, 2024
Training in progress, step 4020
dcd2363
verified
Audreygyj
commited on
Dec 8, 2024
Training in progress, step 3216
8a74704
verified
Audreygyj
commited on
Dec 8, 2024
Training in progress, step 2412
0f305a4
verified
Audreygyj
commited on
Dec 8, 2024
Training in progress, step 1608
4b73ef9
verified
Audreygyj
commited on
Dec 8, 2024
Training in progress, step 804
f59bca7
verified
Audreygyj
commited on
Dec 8, 2024
initial commit
25b6a06
verified
Audreygyj
commited on
Dec 7, 2024