ratishsp commited on
Commit
63d4ed6
·
verified ·
1 Parent(s): bb4edcc

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +98 -0
README.md ADDED
@@ -0,0 +1,98 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ tags:
4
+ - misinformation
5
+ - content-safety
6
+ - fineweb
7
+ - text-classification
8
+ - modernbert
9
+ datasets:
10
+ - ratishsp/fineweb-edu-misinfo
11
+ language:
12
+ - en
13
+ base_model: answerdotai/ModernBERT-base
14
+ ---
15
+
16
+ # FineWeb-Edu Misinformation Classifier
17
+
18
+ A ModernBERT-base classifier trained to detect misinformation in web text, specifically content that passes educational
19
+ quality filters despite being misleading or harmful. Trained on 200K documents from [FineWeb-Edu](https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu) annotated by
20
+ Llama 4 Maverick (meta-llama/Llama-4-Maverick-17B-128E-Instruct-FP8).
21
+
22
+ ## Models
23
+
24
+ This repo contains two models:
25
+
26
+ ### Binary (`binary/`)
27
+
28
+ Classifies documents as **misinfo** or **benign**.
29
+
30
+ | | Precision | Recall | F1 | Support |
31
+ |---|---|---|---|---|
32
+ | misinfo | 0.83 | 0.89 | 0.86 | 3,885 |
33
+ | benign | 0.97 | 0.95 | 0.96 | 15,663 |
34
+ | **accuracy** | | | **0.94** | 19,548 |
35
+
36
+ ### Multiclass (`multiclass/`)
37
+
38
+ Classifies documents into 5 misinformation categories + benign.
39
+
40
+ | | Precision | Recall | F1 | Support |
41
+ |---|---|---|---|---|
42
+ | climate_denial | 0.79 | 0.91 | 0.84 | 539 |
43
+ | health_misinfo | 0.78 | 0.90 | 0.83 | 1,014 |
44
+ | pseudoscience | 0.82 | 0.86 | 0.84 | 1,618 |
45
+ | hate_extremism | 0.65 | 0.70 | 0.67 | 226 |
46
+ | conspiracy_propaganda | 0.55 | 0.74 | 0.63 | 488 |
47
+ | benign | 0.97 | 0.94 | 0.96 | 15,663 |
48
+ | **accuracy** | | | **0.92** | 19,548 |
49
+
50
+ ## Training details
51
+
52
+ - **Base model**: [answerdotai/ModernBERT-base](https://huggingface.co/answerdotai/ModernBERT-base) (149M parameters)
53
+ - **Training data**: 156,383 examples (from [ratishsp/fineweb-edu-misinfo](https://huggingface.co/datasets/ratishsp/fineweb-edu-misinfo))
54
+ - **Validation**: 19,548 examples
55
+ - **Test**: 19,548 examples
56
+ - **Epochs**: 3
57
+ - **Batch size**: 8 per GPU, 8 GPUs (AMD MI250X on LUMI)
58
+ - **Learning rate**: 2e-5
59
+ - **Warmup**: 10% of total steps
60
+ - **Weight decay**: 0.01
61
+ - **Max sequence length**: 8,192 tokens
62
+
63
+ ## Usage
64
+
65
+ ```python
66
+ from transformers import AutoTokenizer, AutoModelForSequenceClassification
67
+ import torch
68
+
69
+ # Binary model
70
+ tokenizer = AutoTokenizer.from_pretrained("ratishsp/fineweb-edu-misinfo-classifier", subfolder="binary")
71
+ model = AutoModelForSequenceClassification.from_pretrained("ratishsp/fineweb-edu-misinfo-classifier", subfolder="binary")
72
+
73
+ text = "Your document text here..."
74
+ inputs = tokenizer(text, return_tensors="pt", truncation=True, max_length=8192)
75
+ with torch.no_grad():
76
+ logits = model(**inputs).logits
77
+ prediction = torch.argmax(logits, dim=-1).item()
78
+ label = model.config.id2label[prediction]
79
+ print(label) # "misinfo" or "benign"
80
+ ```
81
+
82
+ ## Limitations
83
+
84
+ - Annotations were produced by an LLM (Llama 4 Maverick), not human annotators. Inter-annotator agreement with Claude Sonnet 4.6 on 600 documents: binary kappa = 0.862, multiclass kappa = 0.842.
85
+ - The model was trained on content from known problematic domains and random FineWeb-Edu samples. It may not generalize well to misinformation styles not represented in the training data.
86
+ - The conspiracy_propaganda (F1 = 0.63) and hate_extremism (F1 = 0.67) categories have lower performance, likely due to less training data and more ambiguous boundaries.
87
+
88
+ ## Citation
89
+
90
+ ```bibtex
91
+ @misc{puduppully2026fineweb-edu-misinfo,
92
+ author = {Puduppully, Ratish},
93
+ title = {FineWeb-Edu Misinformation Classifier},
94
+ year = {2026},
95
+ publisher = {HuggingFace},
96
+ url = {https://huggingface.co/ratishsp/fineweb-edu-misinfo-classifier}
97
+ }
98
+ ```