Is this model supported for finetuning with flash attention ?
#4 opened 5 months ago
by
thaodd11
MMLU Performance After Token Training
👍
2
#3 opened about 1 year ago
by
adol01