fix(config): passing gradient_checkpoint_kwargs (#1412) b1e3e1b unverified Nanobit commited on Mar 19
Feat(readme): Add instructions for Google GPU VM instances (#1410) 868c339 unverified Nanobit commited on Mar 16
Remove unsupported python version 3.9 from README (#1364) [skip ci] 3765747 unverified Nicolas Rojas commited on Mar 6
ADD: push checkpoints to mlflow artifact registry (#1295) [skip ci] d756534 unverified JohanWork Nanobit winglian commited on Feb 26
chore: update readme to be more clear (#1326) [skip ci] c6b01e0 unverified Nanobit commited on Feb 26
fix(readme): Clarify doc for tokenizer_config (#1323) [skip ci] 2ed52bd unverified Nanobit commited on Feb 24
fix(readme): update inference md link (#1311) [skip ci] 3d2cd80 unverified Nanobit commited on Feb 21
Scheduler implementation of Continual Pre-Training of Large Language Models: How to (re)warm your model? (#1273) 8430db2 unverified jinwonkim93 commited on Feb 13
allow the optimizer prune ratio for ReLoRA to be configurable (#1287) 4b997c3 unverified winglian commited on Feb 12
add contact info for dedicated support for axolotl [skip ci] (#1243) dfd1885 unverified winglian commited on Feb 1
Feat/chatml add system message (#1117) 98b4762 unverified mhenrichsen Mads Henrichsen winglian commited on Jan 25
Fine-Tuning Mistral-7b for Real-World Chatbot Applications Using Axolotl (Lora used) (#1155) cc25039 unverified Tilemachos Chatzipapas twenty8th winglian commited on Jan 23
set fp16 to false if bf16, update bf16: auto in example YAMLs (#1122) [skip ci] 782b6a4 unverified winglian Nanobit commited on Jan 22
feat(dataset): add config to keep processed dataset in memory (#1152) 3db5f2f unverified Nanobit commited on Jan 20
Add shifted sparse attention (#973) [skip-ci] 1d70f24 unverified jrc joecummings winglian commited on Jan 18
Agnostic cloud gpu docker image and Jupyter lab (#1097) ece0211 unverified winglian commited on Jan 16
fix(readme): clarify custom user prompt [no-ci] (#1124) 9cd27b2 unverified Nanobit commited on Jan 16
Add: mlflow for experiment tracking (#1059) [skip ci] 090c24d unverified Johan Hansson winglian commited on Jan 9
Cosine learning rate schedule - minimum learning rate (#1062) 04b978b unverified ricdomolm winglian commited on Jan 9
feature: better device mapping for large models (#918) bdfefaf unverified dg-kalle Karl-Johan Alm winglian commited on Jan 5