Göktuğ Düşünen PRO
AI & ML interests
Head of AI @ Werea · AI/ML Engineer
AI/ML Engineer building end-to-end intelligent systems across LLMs, NLP, Computer Vision, RAG, Retrieval, Agents, and Applied AI.
From data and model development to fine-tuning, evaluation, optimization, deployment, and production AI systems.
Recent Activity
posted an update 1 day ago
🇹🇷 We started with one question:
**How much of the Turkish AI stack can we build openly?**
Today, Werea has grown to **19 open models on Hugging Face.**
Not just LLMs.
📄 Document AI — Werea-DocOCR-1B
🛡️ Cybersecurity — Werea-NanoSOC-8B
🔐 Privacy / KVKK — Werea-KVKK-Agent-4B
🔍 Retrieval — DUSUNEN-Rota-270M
🎙️ Speech — Werea-TSS
🧠 Turkish NLP — NER, NLI, QA, Intent, Sentiment, Topic & more
👁️ Computer Vision — Gemstone
And we want the results to be measurable.
Some of our published benchmarks:
🏷️ NER → **91.7% F1** — WikiANN-tr
🗂️ Topic → **92.8% accuracy** — TTC4900
🎯 Intent → **88.2% accuracy** — MASSIVE-tr
🧩 NLI → **74.5% accuracy** — XNLI-tr
❓ QA → **72.7% F1** — TQuAD2
📄 DocOCR → **0.15% CER** on our held-out Turkish enterprise document test
Our goal isn't to upload as many models as possible.
Our goal is to build an **open Turkish AI ecosystem**:
Models.
Datasets.
Benchmarks.
Demos.
Real applications.
Built from Türkiye. 🇹🇷
Open to everyone.
🤗 Explore Werea:
https://huggingface.co/Werea-co
📄 Document AI:
https://huggingface.co/Werea-co/Werea-DocOCR-1B
🛡️ NanoSOC:
https://huggingface.co/Werea-co/Werea-NanoSOC-8B
🔐 KVKK Agent:
https://huggingface.co/Werea-co/Werea-KVKK-Agent-4B
🔍 Rota:
https://huggingface.co/Werea-co/DUSUNEN-Rota-270M-v3
If you're building Turkish AI, follow Werea — there's much more coming.
posted an update 5 days ago
🇹🇷 We trained a 1B OCR model specifically for Turkish enterprise documents.
**Werea-DocOCR-1B v2**
The result surprised us:
LightOnOCR-2 base → **64.2% CER**
Werea-DocOCR v1 → **~8.1% CER**
Werea-DocOCR v2 → **0.15% CER** 🚀
Evaluated on a held-out 72-page test set across 12 Turkish document types and 3 different capture conditions.
📄 12 Turkish enterprise document types
🧪 12,960 synthetic training pages
📱 Digital + scanned + phone photos
📊 Tables → structured Markdown
⚙️ Full-parameter fine-tuning
🖥️ Trained on a single RTX 3090
It handles:
• e-Invoices
• rental contracts
• bank receipts
• payroll documents
• insurance policies
• vehicle documents
• official correspondence
• trade registry documents
• SGK-style tables
• and more.
**Model 🤗**
https://huggingface.co/Werea-co/Werea-DocOCR-1B
**Dataset 📚**
https://huggingface.co/datasets/Werea-co/werea-tr-doc-ocr-enterprise-v2
**Werea 🇹🇷**
https://huggingface.co/Werea-co
We're building open AI models from Türkiye.
This is just the beginning.
#HuggingFace #OCR #DocumentAI #TurkishAI #OpenSourceAI #ComputerVision
reacted to jasoncorkill's post with ❤️ 6 days ago
Most public benchmarks collapse model performance into one broad preference signal.
That makes it hard to understand which capabilities differentiate between models. It's also almost impossible to inspect the evidence behind it. So @RapidataAI is releasing Benchmark.AI.
We started with an SVG generation benchmark including 42 models, 500 prompts, 1.9M+ human judgements, 300K+ match-ups.
We evaluate models separately on Preference, Alignment and Coherence, while making the prompts, outputs, match-ups and methodology public.
Full dataset: https://huggingface.co/datasets/Rapidata/svg-benchmark
Full benchmark: https://www.benchmark.ai/svg
Methodology feedback and benchmark suggestions very welcome!