File size: 4,774 Bytes
b296ad4
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
# TinyQuery: how the model works

TinyQuery is a 139,738,113-parameter decoder trained from random weights for a narrow task: turn a question plus database schema and tool definitions into one JSON action. It uses established transformer components and a learned source-copy head. It is an experiment, not a claim of a new research architecture.

```mermaid
flowchart LR
    A[Schema, tools and question] --> B[Byte BPE: 4,082 tokens]
    B --> C[12 causal decoder blocks]
    C --> D[Generate next-token probabilities]
    C --> E[Copy probabilities over input tokens]
    C --> F[Learned sigmoid gate]
    D --> G[Mixed next-token distribution]
    E --> G
    F --> G
    G --> H[Raw JSON action]
    H --> I[Validate tool name and arguments]
    I --> J[MCP tools/call payload]
```

The model sees everything needed for a particular request in its prompt. It must learn the language of the request, the required SQL operation, the correct runtime tool, and how to preserve identifiers and literal values. The inference CLI prints the action; it does not connect to a database or execute the tool.

## Decoder

| Component | Configuration |
|---|---|
| Tokenizer | Byte BPE trained only on training text; 4,082 tokens |
| Hidden width | 1,024 |
| Decoder layers | 12 |
| Attention | 16 query heads, 4 key/value heads, head width 64 |
| Positions | Rotary position embeddings, theta 10,000 |
| Normalization | RMSNorm before attention and feed-forward layers |
| Feed-forward | SwiGLU, intermediate width 2,816 |
| Output projection | Tied to the input token embedding |
| Context limit | 2,048 tokens; training coverage is shorter and reported with the dataset |
| Training-only auxiliary head | Three action classes: call, clarify, answer |

Attention is causal: a position can use earlier tokens but cannot see future answer tokens. Grouped-query attention shares four sets of keys and values across sixteen query heads, reducing the KV cache used during streaming. The cache preserves earlier attention keys and values so generation does not recompute the entire prompt for each token.

## Learned copying

Ordinary next-token generation initially memorized familiar table and tool names. Dataset variants and a copy head address this specific failure. The copy head projects decoder states into 128-dimensional queries and keys, attends to the supplied context and question, and adds together attention mass for repeated occurrences of the same token.

For each next token, a learned gate combines two distributions:

`P(token) = gate × P_generate(token) + (1 − gate) × P_copy_from_input(token)`

The copy distribution excludes generated answer text. This prevents the model from repeatedly copying its own earlier mistakes. The gate and attention are learned; there is no hard-coded replacement of predicted tool names, no SQL template compiler, no constrained decoding, and no hidden teacher or retrieval fallback in neural evaluation.

This adapts the established [pointer-generator idea](https://aclanthology.org/P17-1099/). Data and architecture changed together during development, so improvements are not presented as a controlled architecture ablation.

## Training and output

Early stages give answer tokens weight 1 and prompt tokens weight 0.15, plus 0.05 times an auxiliary action-classification loss. A late refinement uses prompt weight 0, focusing optimization on answers; final checkpoint selection can retain an earlier candidate if that refinement does not improve development results. The prompt contribution helps learn vocabulary from scratch; the answer contribution teaches the requested behavior. The final curriculum also samples particular context variants more often to counter identifier-format memorization.

The teacher supplies language variations and a verification pass. Programmatic recipes supply reference SQL and tool actions. Student weights originate entirely from random initialization in this experiment; later stages continue those same weights and add randomly initialized copy parameters.

The output contract is either a call, a clarification, or a short answer:

```json
{"action":"call","name":"execute_sql","arguments":{"query":"SELECT name FROM customers WHERE city = 'Delhi';"}}
```

The adapter checks the action against the supplied tool schema and can serialize a JSON-RPC `tools/call` request. JSON validity, correct tool selection, exact project scope and SQL result correctness are separate evaluation measures. A valid JSON action alone does not establish a correct query.

The model's scope is bounded SQL and tool use in four language styles. It is not a general conversational model or a database knowledge store. Final measured results, training time and limitations belong in the release model card.