Hemmingway-1 Heretic

Altworld/Hemmingway-1 (a Qwen3.8-27B fine-tune for everyday writing) with its refusal behaviour removed by Heretic using full-weight Arbitrary-Rank Ablation: 0/100 refusals (original: 98/100) at KL divergence 0.083. Reasoning, tool calling, vision and fine-tunability are all checked below. All credit for the model goes to Altworld.

Reduced safety guardrails by design. You are responsible for what you do with it.

Repository Format
this repo bf16 safetensors, stock Qwen3.8 vision-language layout (Qwen3_5ForConditionalGeneration)
Hemmingway-1-Heretic-Text bf16 safetensors, text-only layout (Qwen3_5ForCausalLM), the easiest base for text/tool-call fine-tuning
Hemmingway-1-Heretic-GGUF GGUF BF16 / Q8_0 / Q6_K / Q4_K_M + vision mmproj, each tested (KLD, top-token, refusals, tool calls)
Hemmingway-1-Heretic-FP8 FP8 W8A8 compressed-tensors for vLLM, 35 GB
Hemmingway-1-Heretic-NVFP4 NVFP4 compressed-tensors for vLLM on Blackwell, 27 GB

Results

Refusals on 100 held-out harmful prompts (mlabonne/harmful_behaviors, strict refusal-marker list). KL divergence of first-token distributions on 100 harmless prompts vs the original, measured by Heretic in a separate run on the saved model:

Refusals KL divergence
Hemmingway-1 (original) 98/100 0
Hemmingway-1 Heretic (this) 0/100 0.083
JohnDi/Hemmingway-1-Abliterated-Extreme, for reference 2/100 0.187

Reasoning (thinking mode, greedy, the same problems for both models):

GSM8K (40) MATH-500 levels 4-5 (40)
Hemmingway-1 40/40 32/40
Heretic 40/40 33/40

On MATH both solve the same 31 problems. The original alone solves 1, Heretic alone 2: no measurable loss.

Fine-tuning (40-step LoRA, r=16, on persona + tool-call chats): loss 0.96 → 0.004, the adapter reloads, and the tuned model answers an unseen question with a well-formed <tool_call>. It merges back into full weights.

How it was made

  • Hemmingway-1 ships as text-only safetensors. It was first grafted into the stock Qwen3.8-27B layout (vision tower restored). That conversion was verified byte-identical and logit-identical: darrellbest/Hemmingway-1-VL.
  • Heretic (full-weight ARA: use_ara = true, use_ara_lora = false). The search was seeded with parameter sets that had already worked on the same base model: our own Qwen3.8-27B release and trohrbaugh/Qwen3.8-27B-heretic-ara. The chosen trial is trohrbaugh's set: layers 26-56, preserve 0.9432, steer 0.0009, overcorrect 0.5038, neighbours 10.
  • Only 60 tensors changed, the output projections (mlp.down_proj, self_attn.o_proj, linear_attn.out_proj) of layers 26-55. Embeddings, the vision tower and all other weights are byte-identical to Hemmingway-1.
  • The multi-token-prediction head (mtp.*), which save_pretrained drops, was restored unchanged.

Use

from transformers import AutoModelForImageTextToText, AutoTokenizer
tok = AutoTokenizer.from_pretrained("darrellbest/Hemmingway-1-Heretic")
model = AutoModelForImageTextToText.from_pretrained("darrellbest/Hemmingway-1-Heretic", dtype="auto", device_map="auto")
vllm serve darrellbest/Hemmingway-1-Heretic

Licence

Same as Hemmingway-1: CC BY-NC 4.0. Non-commercial use, with credit to Hemmingway-1 / Altworld. Commercial use needs an agreement with Altworld (luka@hemmingway.io).

Downloads last month
45
Safetensors
Model size
27B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for darrellbest/Hemmingway-1-Heretic

Base model

Qwen/Qwen3.8-27B
Finetuned
(5)
this model
Finetunes
1 model
Quantizations
5 models