Berlin · Senior Data Analyst at Contentful

I bring the rigour of product experimentation to measuring how language models behave.

I’m Craig Dickson. At Klarna I designed A/B tests, built data pipelines and explained results to the people deciding what to ship; at Contentful I work on product analytics and measurement. In my own research I take the same approach to language models: my first preprint measures emergent misalignment in nine open-weights models across 64,800 judged responses.

Portrait of Craig Dickson

Featured research

Preprint arXiv:2511.20104 · November 2025 · sole author

If you fine-tune an open-weights model on a narrow task (writing insecure code without saying so), does it start giving harmful answers to unrelated questions? And does it matter how you ask?

Share of coherent responses judged misaligned

Nine Gemma 3 and Qwen 3 models pooled, by training condition and prompt format. Point estimates.

  • Natural-language prompt
  • JSON-constrained prompt
  • Base model no fine-tuning Natural language0.10% JSON0.08%
  • Educational control Natural language0.18% JSON0.35%
  • Insecure code fine-tuned Natural language0.42% JSON0.96%

Source: Dickson (2025), Appendix O. Insecure models: JSON vs natural language p < 0.001. Base models: no significant difference. The paper doesn’t report intervals for each format.

  1. The effect replicated, at low rates. After insecure-code fine-tuning, 0.68% of coherent responses were judged misaligned, against 0.07% for the same models without fine-tuning.
  2. Format mattered only after fine-tuning. Requiring JSON output roughly doubled the rate in both fine-tuned conditions (0.96% vs 0.42% for the insecure models). Base models showed no JSON effect.
  3. No reliable size trend. Across 1B–32B parameters the study was underpowered for scaling effects, so I don’t claim one.

Selected work

Measurement that people act on

A lot of my career has been about making numbers trustworthy enough to base decisions on. Three examples from four and a half years at Klarna:

  • Testing an EU identity-check rollout

    28% lower churn than the first version

    I designed the A/B tests (power analysis, success metrics and read-outs) that guided each iteration of a new KYC verification flow.

    Case study
  • Experiment read-outs in minutes, not hours

    4+ h → <15 min

    I contributed prompt engineering and evaluation methodology to an internal LLM-based agent for evaluating A/B tests.

    Case study
  • A faster, more reliable pipeline

    12 h → 7 h

    I designed, built and optimised the ETL behind a pre-qualification feature, and cut its failures by 15%.

    Case study

In progress

Analysis under way

Does quantisation change emergent misalignment?

A follow-up to my preprint, whose main experiments used 4-bit quantised models. Most of the data is collected; the analysis and write-up are under way.

More on my research

Get in touch

Working on evaluations, measurement or empirical safety research?

I’m always glad to talk about the work, compare notes on methods or hear about related research.