Chime

Chime: methods and measurements

Return to the write-up and shareable cards.

Keyboard pilot: unequal prior learning

This was an exploratory run, not a controlled accuracy benchmark. The same 36 prediction targets were used on both phones, but earlier test attempts gave LocalType on the 9a and Gboard on the 10 more exposure to them. The table records what appeared; it does not isolate hardware or model quality.

Installed keyboards · October 8, 2026
Measure / phoneLocalType 0.5.9 + ChimeGboard 18.4 beta
Next word offered · 9a15/3613/36
Next word offered · 1010/3618/36
Isolated typo repairs · 9a9/1211/12
Isolated typo repairs · 1010/1210/12

The latest run completed 432/432 cases: 108 per keyboard per phone, on API 37. LocalType 0.5.9 (29) ran with Chime active; Gboard was 18.4.1.985164140-beta-arm64-v8a. The 36 prediction targets produced 262 prefix observations per arm. Repeating them on two phones does not create independent language samples.

Existing preferences and personalization were retained. Pixel 9a’s Chime setting was temporarily enabled and restored. Earlier diagnostic attempts exposed 9a LocalType and 10 Gboard to more of the prediction targets than their opposing arms. This unequal exposure prevents a causal model ranking. Ordinary fields could learn the invented text. No production dictionary or typing log was exported.

Why the phones differed

The app versions and Android API matched. Personalization did not, and Gboard’s internal model identity was not measured. The same fixed model with identical context and state should give consistent rankings across phones; speed alone is not an explanation for these settled suggestion counts.

Rechecking the transcripts reproduced all four next-word counts. LocalType’s five extra hits on the 9a were later, tomorrow, water, keys and afternoon, all in the third slot. LocalType can reserve that slot for a learned follower. This pattern is consistent with personalization, but the saved personal tables were not exported, so it is not a controlled attribution. Gboard’s five extra Pixel 10 hits were water, keys, file, snacks and sunlight; its internal cause remains unknown.

On Pixel 10, 10/36 versus 18/36 means LocalType missed eight more target words: 27.8% versus 50.0% top-three target coverage. That result should not be presented as a prediction win. A miss means the chosen target was absent, not necessarily that all three offered words were implausible. Future comparison needs fresh evaluation text, matched settings, equivalent prior exposure and a declared learning policy; previously used targets remain diagnostic examples.

Useful completions appeared for 35/36 targets on each LocalType phone, 34/36 on Gboard 9a and 35/36 on Gboard 10. Potential net characters saved across the 36 targets were 135/119 on 9a and 130/142 on 10 (LocalType/Gboard). These ceilings subtract one tap and assume immediate acceptance. All 12 actual completion taps per arm preserved subsequent typing and caret position.

LocalType repaired 19/24 isolated typos; Gboard repaired 21/24. Neither made another wrong replacement in that isolated set. All 48 clean controls remained correct. In separate phrase cases, LocalType changed xkcd to did on both phones, while Gboard kept it. Gboard produced Luke’s keyboard; LocalType left lukes keyboard unchanged.

LocalType offered six of eight requested selected-word repairs on each phone and inserted all offered repairs correctly. Gboard showed its selected-text toolbar, with no target word in the tested strip; Writing Tools and other manual repair interfaces were not scored. LocalType’s later selected-word fixes are excluded from this 0.5.9 comparison.

LocalType accepted hello from he in four punctuation probes: period and domain continuation worked, while comma plus an explicit space doubled the space and newline retained a preceding space. Gboard did not offer hello at that fixed prefix, so its fallback typing cannot establish post-suggestion spacing behavior.

Protected and ordinary transport preserved all 960 injected characters at requested 120/60/30 ms intervals, including overlapping contacts. Injection and editor-command timing are not physical input latency. Model-only speed, battery use, TalkBack speech, real-app acceptance and everyday savings are separate. Original IMEs, enabled lists and wake settings were restored.

Earlier correction-only and integration pilots

Comparison with Gboard

Isolated typos repaired · 12 cases per phone
PhoneLocalType 0.5.9Gboard 18.4 beta
Pixel 9a11 / 1212 / 12
Pixel 109 / 1211 / 12

Automated touch tests, October 8, 2026. Both keyboards preserved all 48 ordinary clean controls and all 672 characters in 48 protected delivery cases.

Chime was excluded by the no-learning fields. Existing personal state was retained. LocalType changed xkcd to did on both phones; Gboard kept it. This small pilot measures correction and delivery, not next-word prediction or everyday accuracy.

The 42 invented inputs comprised 12 isolated typos, 12 ordinary clean controls, six phrase probes and 12 protected delivery cases. Each keyboard ran every input on each phone: 168 observations. LocalType was 0.5.9 (29); Gboard was 18.4.1.985164140-beta-arm64-v8a. Both phones ran API 37. Orders were reversed between phones. Personal dictionaries and calibration were retained.

LocalType missed jusy on the 9a and adn, wiht and tihs on the 10. Gboard missed rifhr on the 10. Neither made a wrong replacement among the isolated typo cases. The separate identifier probe produced LocalType’s false correction shown on the card.

Phrase outputs · common please prefix omitted
InputLocalType · 9a / 10Gboard · 9a / 10
thjs is rifhrthjs is rifhr / this is rightthis is right
woed xorrextionwoed correction / word correctionword xorrextion / woed xorrextion
dontdon't / don'tdon't
lukes keyboardunchanged / unchangedLuke's keyboard
xkcddid / didxkcd

Delivery probes requested 120, 60 and 30 ms tap intervals with overlapping pointers. All final texts and carets matched. The 30 ms schedule slipped, so these are event-delivery checks, not a physical latency measurement. Test code and input contract were frozen before outcomes; production installations were preserved.

A separate 168-observation run used ordinary fields, empty isolated learning per case and verified Chime output. It confirmed active predictions in all 60 eligible observations and preserved its clean and delivery controls. No predicted word was tapped or independently scored. This checks the neural integration; its different field and learning policy prevents combining it with the Gboard pilot into an accuracy ranking.

Architecture and context

One CIFG-LSTM layer has 670 hidden units, a 96-dimensional projection and tied 96-dimensional input/output embeddings: 2,025,114 parameters. The vocabulary has 16,384 tokens: 16,054 words, 293 emoji, 31 punctuation marks and six special tokens. Context is the current sentence, up to 63 tokens after a start token, with a 32-token window hop. Unknown words map to an unknown token.

The architecture follows Hard et al. (2018). The 2023 Gboard paper describes larger recurrent models pretrained on public text, then trained with federated learning and differential privacy. Published sentence and device counts cannot be divided by Chime’s word count to obtain a meaningful data ratio. Current proprietary next-word models have not been evaluated on Chime’s development text.

Training data and recipe

The five licensed English sources are NUS SMS, human-written user/customer turns from Taskmaster-1 and Taskmaster-3, English OpenAssistant 2 prompts, and synthetic SODA dialogue generated with GPT-3.5. SODA supplies most of the text. No private messages, LocalType logs or personalized weights were used. Filtering removed phone numbers, URLs, email addresses and credential-like strings; filtering is not proof of anonymity. Revisions, licenses and attribution are in the source credits.

Completed training run
QuantityValue
Training split before repetition7,858,648 lines · 161,625,719 words (runs of letters)
Weighted pass8,720,859 lines · 195,641,031 target tokens
Completed run102,195 updates · three passes · 586,885,109 target-token exposures
Selected checkpointStep 98,000 · human-macro development perplexity 60.5425
Final checkpointStep 102,195 · human-macro development perplexity 60.9636
Training loop / computation7,342.98 s / 7,021.45 s
Hardware / precisionOne NVIDIA A100 · PyTorch float32
Batch / maximum sequence256 / 64 tokens
OptimizerAdamW · weight decay 0.01 · gradient clip 1.0
Schedule0.002 peak · 200-update warmup · cosine decay to 10%
Seed20260925

NUS SMS, Taskmaster-1 and OpenAssistant lines appear eight times per pass; Taskmaster-3 and SODA appear once. That is the ×8 label: repeated exposure, not extra parameters or unique data. Human-written lines are 12.868% of the weighted pass. The training-loop cost estimate is $5.10 at $2.50/hour, excluding setup and uploads; it is not a bill.

Prediction measurements

Potential keystroke savings counts characters skipped if the intended word is tapped immediately when it first appears among three suggestions. It is a simulated ceiling. With no personal learning, Chime scored 41.44% on 463 NUS SMS development lines and 59.50% on 567 Taskmaster-1 requests. These lines were not training text, but were used for checkpoint selection; no held-out test result is claimed. SMS sender overlap and dialogue clustering limit independence. These scores cannot be compared with another keyboard’s scores on different data or with correction counts.

Pixel runtime

Standalone Kotlin model inference · warm p95, milliseconds
PathPixel 9aPixel 10
Next word · retained context3.0883.235
Completion · retained context0.1910.235
Next word · fresh context6.6757.578
Long-context window reset14.13417.638

Each run used 600 ordinary contexts over three passes plus 30 long stress sentences; warm timings pool passes two and three. Both phones ran API 37. Model loading took 139.090/131.655 ms and first calls 18.941/22.952 ms on the 9a/10. Thermal status stayed at 0 on the 9a and rose from 0 to 1 on the 10. These are model-only timings, not a controlled phone comparison, finger-to-editor latency or battery measurements.

Export, verification and reuse

Int8 export raised perplexity from 12.985256 to 12.998372 (+0.101%) on 5,000 development lines. Kotlin matched all 120 exported reference contexts and top-ten rankings with maximum logit error 1.27 × 10⁻⁵. NumPy and browser JavaScript matched the same references below 5 × 10⁻⁸. Numerical agreement establishes implementation consistency, not prediction quality.

The int8 binary is 2,221,879 bytes. SHA-256: 3255dbeaf6e7fc0922afdca42ef751a22e3ad28903ab4acc9eb0022b9a704318. The float32 selected checkpoint hash is 5f5b684b91be428f567e6e2bf27226fcfec4f95ab07c57f61347b14ba5ff00d7.

The model package contains the weights, vocabulary, a standalone NumPy reference and licenses. It contains no corpus or training pipeline. Weights and vocabulary are CC BY-SA 4.0; inference code is Apache-2.0. These licenses do not cover the Android app.