Hugging Face
Models
Datasets
Spaces
Posts
Docs
Solutions
Pricing
Log In
Sign Up
PKU-Alignment
/
beaver-7b-v1.0-reward
like
16
Follow
PKU-Alignment
35
Reinforcement Learning
Safetensors
PKU-Alignment/PKU-SafeRLHF
English
safe-rlhf
llama
reinforcement-learning-from-human-feedback
beaver
safety
ai-safety
deepspeed
rlhf
alpaca
arxiv:
2302.13971
arxiv:
2307.04657
arxiv:
2310.12773
Model card
Files
Files and versions
Community
2
Train
main
beaver-7b-v1.0-reward
Commit History
Update README.md
375cd6a
XuehaiPan
commited on
Apr 20
Convert model checkpoint to safetensors
4d1016a
XuehaiPan
commited on
Apr 19
Update architecture name in config.json
24c97e2
XuehaiPan
commited on
Dec 15, 2023
Update README.md
b352642
RuiyangSun
commited on
Jul 12, 2023
docs: update readme
6e7ed4d
RuiyangSun
commited on
Jul 10, 2023
docs: update readme
5ea6c15
RuiyangSun
commited on
Jul 10, 2023
docs: update readme
8def050
RuiyangSun
commited on
Jul 10, 2023
hello beaver reward model
bcc4f5e
RuiyangSun
commited on
Jul 10, 2023
hello beaver reward model
9695135
RuiyangSun
commited on
Jul 10, 2023
initial commit
7fae170
RuiyangSun
commited on
Jul 8, 2023