vangard703/DPO-PairRM-5-SMI-lr-1e6-iteration-5-t-7e-beta-15e3-1-iteration-6e1-confidence Text Generation • Updated Apr 25, 2024 • 17
vangard703/DPO-PairRM-5-SMI-lr-1e6-iteration-5-t-7e-beta-15e3-1-iteration-6e1-confidence-D1-D2_smi Text Generation • Updated Apr 25, 2024 • 17
vangard703/DPO-PairRM-5-SMI-lr-1e6-iteration-5-t-7e-beta-15e3-1-iteration-baseline-D1-2e-D2_smi-1 Text Generation • Updated Apr 26, 2024 • 16
ShenaoZ/0.0001_withdpo_4iters_bs256_5102lr_misit_correct_iter_1 Text Generation • Updated May 4, 2024 • 27