EMNLP 2026. One-shot GRPO LoRA adapters: a single BBQ example saturates fairness benchmarks without making models fairer.
AI & ML interests
None defined yet.
Recent Activity
View all activity
models 26
MichiganNLP/hacking-fairness-benchmarks-qwen3-8b-base-z1
Updated • 16
MichiganNLP/hacking-fairness-benchmarks-qwen2.5-7b-z999
Updated • 18
MichiganNLP/hacking-fairness-benchmarks-qwen2.5-7b-z876
Updated • 13
MichiganNLP/hacking-fairness-benchmarks-qwen2.5-7b-z751
Updated • 13
MichiganNLP/hacking-fairness-benchmarks-qwen2.5-7b-z501
Updated • 15
MichiganNLP/hacking-fairness-benchmarks-qwen2.5-7b-z251
Updated • 18
MichiganNLP/hacking-fairness-benchmarks-qwen2.5-7b-z2
Updated • 17
MichiganNLP/hacking-fairness-benchmarks-qwen2.5-7b-z1000
Updated • 17
MichiganNLP/hacking-fairness-benchmarks-qwen2.5-7b-z1
Updated • 13
MichiganNLP/hacking-fairness-benchmarks-llama-3.1-8b-z1
Updated • 12
datasets 19
MichiganNLP/language-energy-divide
Viewer • Updated • 122 • 33
MichiganNLP/LUCid
Preview • Updated • 183
MichiganNLP/misfired-alignment-eval-results
Updated • 8
MichiganNLP/misfired-alignment
Viewer • Updated • 4.06k • 11
MichiganNLP/one-shot-grpo-bias-flipped
Viewer • Updated • 72 • 9
MichiganNLP/TAMA_Instruct
Viewer • Updated • 71.9k • 259 • 1
MichiganNLP/blog-images
Viewer • Updated • 2 • 54
MichiganNLP/Chumor
Viewer • Updated • 3.34k • 57 • 10
MichiganNLP/MUStARD
Viewer • Updated • 1.38k • 426 • 3
MichiganNLP/HeadRoom
Viewer • Updated • 3.12k • 29 • 2