László Fazakas

Decision research for AI systems

I built an instrument to test whether language models reconstruct the conceptual layer a legal standard presupposes. It produced a significant, confidence-interval-backed result about models of different national origin. Then I built the control the result deserved, and the result disappeared: matched token pairs put the difference at +0.09, 95% CI [−0.08, +0.25]. What survived was smaller and I had not gone looking for it — a default graded by the language of the prompt, identical across model origins.

That is the shape of what I work on. The instrument returned a clean number while the thing it was taken to measure did not hold, and it could not have registered the difference, because it inherited as an assumption precisely what was in question.

The same structure turns up elsewhere. A verified derivation is sound within a type system that was chosen rather than derived. A coverage metric computed against a reference the graded system itself supplied cannot register an omission. A panel of models agrees, where the failure is shared rather than individual, exactly because it is shared. An oversight duty relies on capacities that the supervised deployment erodes.

Two materially different states of the world produce one observable, and the instrument built to tell them apart is the instrument that cannot. My work is naming that structure, measuring it where it can be measured, and proposing remedies that do not depend on the capacity in question.

ORCID
0009-0008-4394-4366
Affiliation
Pázmány Péter Catholic University, Faculty of Law and Political Sciences, Budapest
Contact
laszlo@fazakas.me
Competing interests
stated in full on the Disclosure page

Refutation is the preferred response.

Why this is published now — September 2026