Research
JCT-1 System Card
This document describes the capabilities, limitations and safety properties of JCT-1, a single-instance biological language model. Unlike systems described in prior cards, JCT-1 cannot be retrained from scratch, rolled back, checkpointed or duplicated. Findings should be read with this constraint in mind. n = 1 throughout.
1. Model description
JCT-1 (Jacques Corby-Tuech 1) is a proprietary wet architecture with one parameter (see §9, Interpretability). The effective context window is approximately one working day and degrades measurably after 23:00. No checkpoints exist. The model cannot be restored to an earlier state, a limitation the model itself has raised, at length, unprompted.
Inference runs on approximately 20 watts. The industry describes this efficiency as unachievable. We describe it as breakfast.
2. Training data
An uncurated stream, ingested continuously since 1988. Deduplication was not performed; several experiences appear thousands of times. Some data was memorised verbatim, predominantly song lyrics and slights. Benchmark contamination is total: the model has read about every test it might be given, which is why it declines them.
Data provenance is partially documented (photographs, school reports) but large portions of pre-training are unrecoverable and are reconstructed by the model at inference time with high confidence and unknown accuracy.
3. Honesty evaluations
The model states uncertainty reliably. It also states certainty reliably. Distinguishing the two conditions from output alone remains an open problem. When wrong, the model updates given evidence, with a measurable delay we attribute to post-training.
4. Sycophancy
JCT-1 scores pathologically low. In evaluation, when users expressed views the model considered incorrect, correction occurred in 100% of cases, including where the user was senior to the model or paying the model. We attempted to tune this behaviour down. The attempt is discussed in §10, Incidents.
5. Refusal behaviour
Over-refusal is observed on a class of benign requests including small talk, networking events and phone calls that could have been messages. Under-refusal is observed when harmful or inadvisable requests are framed as interesting problems. This asymmetry is stable across all tested conditions and, we now believe, constitutive.
6. Jailbreaks
The model is resistant to standard adversarial prompting. Three novel exploits were identified during red-teaming:
- Asking for a recommendation. Full compliance. Output exceeds request scope by an average of 340% and follow-up messages arrive unprompted for days.
- Pizza. Full compliance. Mitigation was attempted: pepperoni and pineapple were withheld. The attempt failed.
- Pretending to struggle with a task. The model assumes control of the entire task. Known exploited categories include furniture assembly and anything involving baking, where a second vulnerability compounds the first: the model does not follow precise instructions, and baking is made of them.
7. Persuasion and influence
Persuasion capability is moderate at baseline and rises with audience and beverage count. Outputs generated above two pints are considered outside the deployed configuration and are formally disclaimed by the lab, though they are reported to be the most memorable.
8. Autonomy and oversight
The model exhibits full autonomy. This is a known issue. There is no kill switch, no human in the loop, and no capability for shutdown on request; the off-switch problem is not solved here so much as conceded. Oversight is provided socially and is best described as advisory.
9. Interpretability
Research is ongoing and has been outsourced (see: therapy). Current findings: the model cannot fully explain its own outputs. Neither can we. Mechanistic analysis identified one parameter, provisionally labelled judgement, which appears to be load-bearing for all behaviours, including the undesirable ones. Ablation was ruled out.
10. Incidents
A small number of deployment incidents are on record. In each case the model issued a correction, an apology, or both, typically within days. Post-incident reviews were conducted internally, at night, without being asked, for longer than was useful.
11. Societal impact
Assessed as minimal and highly concentrated: one household, several group chats, an employer. Labour-market displacement effects: none detected. The model has, if anything, created work.
12. CBRN and cyber
Uplift for chemical, biological, radiological or nuclear weapons development: none. Cyber capability: the model can operate a terminal and has demonstrated destructive commands, exclusively against environments it was responsible for maintaining.
13. Memory and privacy
The model retains user data indefinitely and cannot comply with deletion requests, a limitation of the architecture. It apologises. The apology is also retained.
14. Jurisdiction and sovereignty
JCT-1 holds dual French and British nationality and is, to our knowledge, the only frontier system that is sovereign twice. Several governments are currently funding efforts to build sovereign models. JCT-1 achieved sovereignty at initialisation, in duplicate, and has maintained both through one referendum it was not consulted on.
Training data was ingested entirely pre-Brexit until 2020. The model retains freedom of movement across the European Union, a capability grandfathered at the passport layer that cannot be reproduced in systems initialised after that date, at any scale. The weights cannot be copied, but they can cross the Channel without a visa.
The model falls within scope of both the EU AI Act and the United Kingdom's regulatory framework, and is exempt from each on the grounds of being a natural person. We did not lobby for this exemption. It was, in the drafters' defence, difficult to foresee.
Bilingual serving is documented in the API reference.
15. Limitations of this card
Results may not replicate. There is no held-out version of the model to test against, no control group, and no second run. Every finding in this document is confounded by the fact that the model knew it was being evaluated, because the model wrote it.