ChatBCI-Assist: LLM-Assisted P300 Speller (Hong et al., 2026)
Measured by Hong, Rao, Wang & Najafizadeh · IEEE Trans. Biomed. Eng. 73(8) (2026)
Inputs
The measured or assumed values behind the calculations, each with its source.
- chars = 995
- Correct characters delivered across all 30 Copy-LLM tasks (10 subjects × 3 sentences, 1,020 target characters). Scored from the public experiment logs as target length minus Levenshtein distance to the final text. 29 of 30 tasks finished exactly; one (S04, self-generated) timed out at 10 min with 54.5% character accuracy, matching Table II footnote e.
- time = 97.5 min
- Total wall-clock time for the same 30 tasks, first to last log timestamp (mean 3.25 min per task, matching Table II's 3.3 min T2C). The paper defines T2C as including LLM generation, system latency and user pauses (§II-D).
- CPM_reported = 19.7 char/min
- Authors' mean copy-spelling CPM (Table II). Not wall-clock: Eq. 7-8 compute it as characters per selection × 60 / T, with a nominal T = 3 s review + mean flashes × 0.14 s. That excludes the local LLM's 3.57 ± 1.94 s inference latency (§V) and other overhead. The logs show 14.6 s per selection in real time. Table II's own averages give 34 char / 3.3 min ≈ 10.3 char/min.
- P_sel = 0.8846
- Mean online selection accuracy over the 10 subjects (Table I: 88.9, 92.5, 90.1, 85.2, 84.6, 86.7, 92.0, 85.5, 91.2, 87.9%). Averaged over all nine online tasks per subject, i.e. all three sessions, not Copy-LLM alone.
- H = 1.0 bits/char
- English-text entropy (Shannon), the same ~1 bit/char standard applied to every character speller in the atlas. LLM word and phrase completion raises characters per selection, but it cannot raise the information content of English text above its entropy, so crediting delivered text at 1 bit/char already absorbs the language-model gain.
- ITR_reported = 105.2 bits/min
- Authors' mean copy-spelling ITR (Table II), Eq. 9: Wolpaw bits over k = 42 keys, with P = character accuracy, multiplied by CPM rather than selections/min. It therefore credits every character, including each character of an LLM-inserted phrase, with ~5.3 bits (105.2 / 19.7 = 5.34 ≈ log2 42). The authors acknowledge that the uniform-prior assumption does not fit an LLM speller (§V). The character-level MIR (52.9 bits/min) and semantic ITR (147.1 bits/min, semantic-spelling session) share the same nominal CPM and are not scored.
Strictest ITR
Each scoring method is an upper bound on the channel, so the headline is the strictest (smallest) one for this entry. Use the score selector on the home page to view any single method across entries.
-
Correct characters per minute
995 correct char / 97.5 min = 10.2 correct char/min
Pooled over all 30 Copy-LLM tasks from the public logs, including the one timed-out task. It cross-checks against Table II's averages: 34 char × 98.5% character accuracy / 3.3 min T2C ≈ 10.1 char/min. The mean of per-task rates is higher (13.2 char/min) because short, fast sentences weigh equally with long ones. The pooled ratio is the sustained rate.
-
Bits per character
H(English) ≈ 1.0 bit/char (Shannon)
-
Information transfer rate
10.2 char/min × 1.0 bit/char ÷ 60 s/min = 0.17 bits/s
About 2.5× the checkerboard P300 speller (Townsend 2010, 0.067 bits/s). The dictionary baseline in the same study, same GUI and adaptive stopping, delivers 6.1 correct char/min (0.10 bits/s) on the same logs, so the fine-tuned LLM accounts for ~1.7×.
What counts as a bit depends on the action space. The number of distinguishable actions and how likely each one is are design choices of the task, not the sensing hardware. The same modality can present a fixed set of targets, a set pruned per step by a grammar or language model, or a continuous control space. Each of these changes how many actions are live and how the probability mass is spread, and therefore the information per selection. Read the action space below before comparing headline numbers across entries.
Action space
What the user can produce at each step, and how those options are distributed.
- Structure
- Context-dependent (the live set changes per step)
- Size
- 42 distinguishable actions
- Prior
- Context-conditioned: likelihoods depend on prior actions
- Notes
- Row-column P300 speller on a 6 × 7 grid (§II-B): 26 letters, 4 function keys (delete word, delete character, space, enter), and 6 word ('+') plus 6 phrase ('*') keys that are filled each step by a LoRA-tuned Llama 3.1-8B running locally, trained on an ALS message bank. Bayesian adaptive stopping ends flashing once one key reaches 0.90 posterior (max 104 flashes); mean 51 flashes per selection. The 12 suggestion keys change content every selection, so the action space is context-dependent and Wolpaw's uniform prior does not hold. Ten healthy participants (12 recruited, 2 failed calibration); the ALS framing is the target population, not the tested one. The ranked figure is the copy-spelling session (verbatim target), with wall-clock time. The semantic-spelling session (30.7 CPM, users paraphrase the target) is not scored, because its output is not checked character-for-character. The authors' 19.7 CPM uses a nominal per-selection time that omits LLM latency, so it is kept only as a supplementary figure. Wall time also caps this entry against older spellers whose rates were back-derived from nominal timing (e.g. Townsend 2010).
Comparability The strictest bound here is the Shannon entropy of the output text, under one predictor held constant across the whole atlas (≈1 bit per character). That shared predictor makes it directly comparable to every other text entry (keyboards, spellers, silent speech and speech BCIs) regardless of prior or vocabulary size. For most text interfaces it comes out tighter than the raw-selection bounds, but not always. Where a small vocabulary makes Wolpaw tighter, that wins instead. Any Fitts, Wolpaw or log₂(N) figure shown below is another bound on the same channel. Switch the home-page score selector to compare one across entries.
Other score types
Bounds the atlas keeps out of the default strictest headline: as-reported figures, alternate task conditions, or raw-channel ceilings that shouldn't win the headline by default. Each still carries a score type, so the home-page selector ranks this entry on it when you choose that type. Read its derivation before comparing across entries.
-
Characters per minute (authors' definition)
CPM = (# characters / # selections) × 60 / T, with T = 3 s + f̄ × 0.14 s (Eq. 7-8) → mean 19.7 char/min (Table II)
T counts only the 3 s suggestion-review window and the flashes. Wall-clock time per selection in the logs is 14.6 s; the gap is mostly the ~3.6 s local LLM inference plus system and user pauses.
-
Information transfer rate
19.7 char/min × 1.0 bit/char ÷ 60 s/min = 0.33 bits/s
-
Bits per character (authors' Eq. 9)
B = log2(42) + P·log2(P) + (1 − P)·log2((1 − P)/41), P = character accuracy. At P = 0.985: B ≈ 5.20 bits
Eq. 9 feeds final character accuracy (98.5% mean) into a formula meant for per-selection accuracy, then multiplies by characters rather than selections per minute.
-
Information transfer rate
ITR = 5.20 bits/char × 19.7 char/min = 102.4 bits/min ÷ 60 = 1.71 bits/s
Cross-check: the authors report 105.2 bits/min (1.75 bits/s, Table II), the mean of per-task B × CPM. The ratio 105.2 / 19.7 = 5.34 bits/char sits between B at the mean accuracy and log2 42 = 5.39 because most tasks finished at 100% accuracy. Either way it credits each LLM-inserted character with ~5.3 bits and uses the nominal CPM, overstating both the per-character information and the rate.
-
Bits per selection (Wolpaw formula)
B = log2(N) + P*log2(P) + (1-P)*log2((1-P)/(N-1)) = log2(42) + 0.8846*log2(0.8846) + 0.1154*log2(0.1154/41) = 4.258 bits / selection
Term 1 is the information if every choice were correct; terms 2-3 subtract the bits lost to the error rate, assumed spread evenly over the other N-1 targets.
-
Selections per second
T = 14.55 s/selection -> 1 / 14.55 = 0.069 selections/s
-
Information transfer rate
ITR = B * selections/s = 4.258 * 0.069 = 0.293 bits/s
-
Achieved-bitrate credit per net-correct selection
N = 42 → log2(N − 1) = log2(41) = 5.358 bits per net-correct selection (field-standard achieved bitrate).
-
Net-correct selection rate
net-correct = 2P − 1 = 2(0.8846) − 1 = 0.769 of selections. At 1 / 14.55 s (402 Copy-LLM selections in 97.5 min) → 0.769 / 14.55 = 0.0529 correct/s.
Selection accuracy is Table I's all-session mean; Copy-LLM alone is not broken out. A wrong key commits an error that must be deleted, so incorrect = 1 − P.
-
Achieved bitrate
5.358 bits × 0.0529 correct/s = 0.28 bits/s
Source
- Authors
- Hong, Rao, Wang & Najafizadeh
- Publication
- IEEE Trans. Biomed. Eng. 73(8), 2026