Beyond Turing

Chapter 15 · Part V - The City

Guilt and the Algorithm

Predictive justice, the presumption of innocence and the involuntary conditioning of judgement in the age of risk scores, increasingly computed by AI

In courtrooms, for some years now, a number may appear alongside the case file. A risk score, produced by an information-processing system trained on tens of thousands of past cases, which estimates the probability that a person will commit a new offence. The case file tells a story; the score summarises a population. The case file has a name and a face; the score has a confidence interval. The three disciplines that run through this book converge here as nowhere else: big data, because the score is born of judicial archives that are among the vastest and most delicate repositories of data a society produces; information systems, because contemporary justice is itself a great information system, made of registers, electronic case files and databases that must guarantee integrity, traceability and regulated access; intelligent devices, because the digital body has by now entered the courtroom, from the electronic bracelet that makes enforceable a measure less afflictive than detention, to the data from wearable health and wellness devices that the parties offer as items of evidence. This series has traversed healthcare, business, the body, the mind and even theatres of conflict, always posing the same question: who is the subject? In no other territory does the answer weigh as heavily as here, where what is at stake is neither a commercial offer nor a diagnosis to be confirmed, but the personal liberty and the honour of a human being.

This chapter treats its subject with the same engineering rigour as the preceding ones, but in a more measured register: the delicacy of the matter demands it. And it does so from a precise standpoint, which is best declared at once. Nothing here is on trial: not technology, which is a tool and as a tool must be governed; not the humanity of those who prosecute, those who defend and those who judge, which is the most precious resource of the entire edifice; not the institutions, Italian and worldwide, which are that edifice's pillars. The thesis is another, and it is twofold. First: guilt belongs to the order of judgement, not to the order of calculation, and no statistical refinement can transfer it from the one to the other. Second: human judgement, precisely because it is human, is exposed to involuntary conditioning - affecting everyone, in every direction, nothing excluded - and legal civilisation has been, for centuries, the most refined architecture for the safekeeping of judgement that history knows. The correct engineering question, then, is not whether the algorithm can replace the judge, but whether it aids or disturbs that safekeeping.

1. The risk score as an engineering artefact

From the standpoint of information-systems engineering, a tool for assessing the risk of recidivism is a recognisable pipeline, and it is, in its essence, a big data problem: acquisition of variables (age, prior record, socio-environmental context, answers to structured questionnaires), feature engineering, supervised training on historical outcomes, calibration, delivery of an ordinal score. The five Vs encountered in the chapter The Data and the Subject all appear here, and the most fragile is once again veracity: historical judicial data is an administrative archive, not a mirror of moral reality. Thus far, an engineering artefact like those met in the preceding chapters: it is specified, versioned, validated.

The critical properties emerge, as always in this series, at the threshold, at the point where the number meets the person. The first is statistical in nature: the score is a statement about a class, not about an individual. It asserts that, among people with similar profiles, a certain fraction has had a certain outcome in the past; it asserts nothing necessary about the single case, which remains, in the non-zero measure already encountered in the chapter The Data and the Subject, out of distribution with respect to any training set. On the walls of courtrooms it is written that the law is equal for all, and this is one of the summits of legal civilisation. But equality before the law has never meant the interchangeability of persons: in nature no two blades of grass are identical, and no two defendants have ever been the same. The Gospels guard the same truth with an even bolder image: "even the hairs of your head are all numbered" (Lk 12:7) - numbered because every single one is precious, not because it can be reduced to a number. The score deals with classes of cases; judgement deals with unrepeatable individuals. Statistics seeks the similar; justice guards the unique. Legal equality exists precisely for this: to guarantee that every unrepeatable person receives the same safeguards, not to reduce the person to their statistical class.

The second property is semantic in nature: historical data records proxies, not essences. Archives encode arrests, complaints, convictions - administrative events that depend, in part, on where and how one has looked - and not a person's moral inclination. A model trained on that data learns the geography of enforcement together with the phenomenology of crime, and returns both, indistinguishably, in the form of risk. The third is decisional in nature: every threshold applied to the score generates its own confusion matrix, and in this domain the asymmetry of error - which the chapter The Message and the Target encountered in its extreme form - has a quality that no aggregate metric captures: the false positive is a person subjected to an unjustified restriction; the false negative is a potential victim left unprotected. Neither error is a number, and both demand something the model does not possess: someone who answers for it.

The empirical literature, moreover, counsels sobriety. The international debate opened by journalistic investigations into actuarial recidivism tools (Angwin et al. 2016) spurred the research that then formally demonstrated how the various mathematical definitions of fairness are mutually incompatible when base rates differ across groups (Chouldechova 2017); and controlled studies have documented that the predictive accuracy of one of the most widely used tools proves comparable to that of human assessors with no specific training (Dressel and Farid 2018). This is not an argument for demonising the tool; it is an argument for right-sizing its claims. A well-calibrated score can inform. What it cannot do, by construction, is judge.

There is an exact way to state the difference that runs through this whole chapter, and it is the language of variables. The score is, in the technical sense, a reproducible function: fix the input data, and it returns the same number today, tomorrow and in any courtroom, like a well-written equation. That is its strength, and it is a genuine objectivity - the same objectivity as that of mathematical things, which do not depend on who is looking at them. Human judgement is a function of another kind: it depends on countless variables - training, experience, lived history, the point of view from which each of us looks at reality - and this is why two equally competent panels, two equally honest judges, may weigh the same fact in ways that are not identical. To say this is to accuse no one: it is a description of what judging is, when those who judge are persons and not equations. But the conclusion is not the one the technocrat would expect - replacing the dependent function with the reproducible one - because the two objects are not of the same kind: the score is objective over classes of cases; judgement is answerable for unrepeatable persons. And objectivity without answerability is not yet justice: it is merely arithmetic.

2. Judgement as a human act: involuntary conditioning and the architecture that safeguards it

Here the discourse must become more delicate, and more honest. It would be consoling to set against the algorithm's limits a human judgement immune from limits. It is not so, and cognitive science has documented it for half a century: the human mind, in every profession and at every level of excellence, is exposed to involuntary conditioning. The numerical anchoring described by Tversky and Kahneman (1974) operates in expert assessments as well: experiments conducted with professional judges and jurists have shown that numerical values present in the context - even when avowedly random - can steer, without the decision-maker's knowledge, the quantification of a sentencing demand (Englich, Mussweiler and Strack 2006). The media availability of a case, the sequence in which information is presented, the hindsight that rewrites the predictability of events, the cognitive fatigue of roles that demand hundreds of decisions: these are third-party factors that act on everyone - judges, prosecutors, lawyers, expert witnesses, witnesses, jurors - not through any defect of virtue, but through the very fact of being human. And they act in both directions, nothing excluded: they can incline towards severity as towards leniency, against the accused as in their favour. To acknowledge this is not an indictment of anyone; it is the starting point of every adult epistemology of judging.

There is an engineering image that makes the deep reason for this condition visible. Two memory modules of the same model are identical and interchangeable: replace one, and the system does not notice, because they answer to the same specification and carry nothing with them that the specification does not foresee. There has never existed, and there will never exist, a human being who can be substituted for another. In the one who judges, in the one who prosecutes, in the one who defends and in the one who is judged there are at work, once again, those countless dependent variables which the first section named - personal history, memory, fatigue, the convictions matured over a lifetime - and which no datasheet could ever enumerate. It is the same uniqueness that the preceding section recognised in the defendant: it holds, on the same grounds, for every other actor in the trial. This is why prejudices - everyone's, in every direction - can exert their influence without anyone willing it: it is not a fault, it is the human condition. And this is why a courtroom is not a circuit that processes identical cases, but a meeting of unrepeatable persons which procedure has the task of protecting.

And it is here that the discourse turns around, because legal civilisation has known this since long before cognitive science. The entire edifice of the trial - Italian, European, worldwide - can be read, with an engineer's eyes, as the oldest and most refined architecture of error mitigation ever designed: a fault-tolerant system, a designer would say, built in the knowledge that every single component can fail. Adversarial proceedings oblige every hypothesis to survive its own refutation; the impartiality of the judge separates the one who prosecutes from the one who decides; collegiality and the levels of appeal are deliberate redundancy against single-point failure; the duty to state reasons is traceability of the decision from before engineers invented the phrase audit trail; the prescription to judge beyond all reasonable doubt is a confidence threshold fixed by law to protect the weakest party. In this architecture the defence lawyer is not an obstacle to truth but one of its conditions: the defence function is the correction mechanism that prevents the accusatory hypothesis from verifying itself. It is the lesson Italy delivered to the world with Beccaria, and which the best contemporary doctrine calls garantismo (Ferrajoli 1989): not a naive trust in man, but a mature trust in the procedures man has given himself, knowing himself. The judge who submits to constraints, states reasons, accepts being overturned on appeal, is the highest figure of this awareness: a decision-maker who, unlike any model, knows he can err and has chosen to operate within an architecture that safeguards him.

If this is true, the algorithmic score must be assessed for what it is: the latest arrival among the third-party factors that can condition judgement. The literature on automation bias (Parasuraman and Manzey 2010), already encountered in the medical and military chapters of this series, shows that the output of an automatic system tends to exert a perceived authority greater than its real reliability, precisely because it presents itself in the aseptic garb of a number. A risk score introduced without safeguards does not relieve human judgement of its conditioning: it adds another, powerful and disguised as objectivity. The same architecture that protects the trial from the anchoring of a newspaper headline must therefore protect it from the anchoring of a number in the case file - and that is exactly what the most recent law has begun to do.

3. Digital prejudice: the memory that does not forget and the body that testifies

There is, moreover, a new-generation form of conditioning that acts outside the courtroom, before and after the trial, and which this series is particularly equipped to recognise: digital prejudice. The chapters on the data and on the profile have shown that the record does not forget: what has been written remains queryable, recombinable, liable to resurface. Applied to justice, this technical property produces an effect no legislator had foreseen: the accusation becomes perpetual by architecture. An entry in the register, an arrest, a committal for trial are indexed, shared, commented upon; the eventual acquittal, years later, generates a minute fraction of that visibility. Recommendation systems - which, as seen in the chapter The Customer and the Profile, optimise engagement, not completeness - amplify the emotional peak of the accusation and ignore the silent tail of the acquittal. The result is a penalty foreseen by no code: permanent reputational conviction, executed by an infrastructure that knows neither doubt nor appeal, and which strikes the acquitted innocent with the same efficiency with which it files away the guilty.

Meanwhile, a second stream of data has entered the courtroom through the door of intelligent health and wellness devices. The electronic bracelet makes measures less afflictive than detention technically enforceable, and it is the most mature example of a wearable device placed at the service of proportionality. But more and more often it is the devices born for wellness - the watch that counts steps, the sensor that records the heartbeat, the location - that are offered by the parties as items of evidence: the digital body becomes a witness. Here, transposed, the lesson of the medical chapters of this series applies: just as the signal is not the symptom, so the data is not the proof. A heart-rate trace acquired from a wrist sensor carries with it all the metrological uncertainty encountered in the chapter The Sensor and the Threshold - noise, artefacts, calibration - and becomes evidence only through the scrutiny of adversarial proceedings, expert examination, the judge's assessment: that is, through the human architecture that transforms a number into a procedural fact. The device records; only the trial ascertains.

European law has meanwhile opened a breach against the memory that does not forget, with de-indexing and the right to be forgotten (Court of Justice of the EU, Case C-131/12, 2014; Article 17 of Regulation (EU) 2016/679), recognising that the person does not coincide with the most painful trace the web preserves of them. But the question that concerns this book runs deeper than the norm, and it is the ontological difference the epilogue enumerates: man can repent and change; data cannot. Civil legal orders have always known institutions that give juridical form to this truth - rehabilitation, the extinction of penal effects, the time that restores a possibility - because they presuppose that the human being exceeds every one of their acts, including the worst. The digital profile presupposes the exact opposite: that the person is the sum of their records, frozen in the unhappiest instant of their history. Between these two anthropologies no technical mediation is possible. There is a choice of civilisation, and it must be designed.

4. The presumption of innocence as an ideal system and the rules of the algorithmic age

The presumption of innocence deserves to be looked at with the eyes with which a mathematician looks at an axiomatic system: a construction whose logic is perfect, if perfectly applied. The axiom is stated crisply by the fundamental charters: the defendant is not considered guilty until final conviction (Article 27, second paragraph, of the Italian Constitution); every person charged is presumed innocent until guilt has been legally proven (Article 6, paragraph 2, of the European Convention on Human Rights); and Directive (EU) 2016/343 extended the protection to the public portrayal of the accused, requiring the authorities to refrain from presenting as guilty a person who has not yet been judged. In the ideal system, the theorem that follows is limpid: until the final decision, the person should be considered - and spoken of - as innocent, because every level of appeal exists precisely because the previous one can be overturned.

The gap between the ideal system and its application should be measured not with polemics but with mechanisms - and they are the mechanisms described in the preceding paragraphs: the tempo of the media, which consumes a case in the first forty-eight hours while the trial lives on for years; the architectures of attention, which reward the accusation and not the outcome; digital memory, which makes permanent what the law would wish provisional. Faced with this gap, the legislator of the algorithmic age has drawn boundaries that this series recognises as design invariants. Regulation (EU) 2024/1689 expressly prohibits systems that assess the risk of a person committing an offence on the basis of profiling alone or of personality traits alone (Article 5); it places artificial intelligence systems intended for the administration of justice among the high-risk systems (Annex III), subjecting them to the requirements of risk management, transparency and human oversight (Articles 9 to 15); and with Article 14 it demands that the decision remain effectively in the hands of the one who judges. It is the same lesson as the best-known international leading case on the matter, State v. Loomis (Supreme Court of Wisconsin, 2016): the use of an actuarial score may be admitted only with stringent safeguards, and never as the determining factor of the decision. The framework converges from every direction on a single principle of architecture: the algorithm may inform judgement; it may not, at any point in the pipeline, replace it.

5. The incalculable remainder: guilt is judged, not calculated

At the end of the journey, the deep reason for all these boundaries can be stated simply. Guilt - the real kind, the kind the codes call intent and negligence, imputability and the capacity to understand and to will - presupposes what no training set has ever observed: a free subject, who could have acted otherwise. The score knows correlations between pasts; guilt interrogates a freedom. This is why guilt can be expiated, forgiven, redeemed - and the score cannot. The great traditions that accompany this book converge here with striking clarity: the Jewish teshuvah, the Christian metanoia, the Islamic tawba say, in three different grammars, the same anthropological truth - man never coincides with his act, not even with the worst, because he remains capable of return. A legal order that knows rehabilitation guards, in secular form, the same truth. An infrastructure that freezes the person in their record denies it.

And this is why, in the end, criminal judgement is the place where the difference that gives this book its title shows itself unveiled. The judge has a face and judges a face: responsibility is born, as the reflection on the face already encountered in the chapter The Body and the Sensor teaches, from that encounter which no interface replicates. The defendant who appears before their judge is not a feature vector: they are a human being in the most exposed of conditions, whom the entire architecture of the trial - adversarial proceedings, the defence lawyer, reasonable doubt, the statement of reasons - exists to safeguard. In an information system for justice, then, the dignity of the defendant and the presumption of innocence are not constraints added downstream: they are design invariants, from the first commit. And the system, here more than in any other domain traversed by this series, must know when it does not know - and stop.

The law is equal for all: it is written behind every judge. But no two blades of grass are alike, and no two persons are alike: the equality of the law was invented to protect this difference, not to erase it. Statistics seeks the similar; justice guards the unique. The score can be calculated. Guilt must be judged. They are not, and never will be, the same thing. In a courtroom the most important line of code is the one that falls silent: the one that hands the last word, and the decision, back to a human being who looks another human being in the eyes. There, where the algorithm halts, justice does not end: it begins. Because in the courtroom too - above all in the courtroom - dwells the incalculable remainder.

References cited in the text: Tversky A., Kahneman D. (1974), Judgment under Uncertainty: Heuristics and Biases, Science; Englich B., Mussweiler T., Strack F. (2006), Playing Dice with Criminal Sentences, Personality and Social Psychology Bulletin; Angwin J. et al. (2016), Machine Bias, ProPublica; Chouldechova A. (2017), Fair Prediction with Disparate Impact, Big Data; Dressel J., Farid H. (2018), The Accuracy, Fairness, and Limits of Predicting Recidivism, Science Advances; Parasuraman R., Manzey D. (2010), Complacency and Bias in Human Use of Automation, Human Factors; Beccaria C. (1764), Dei delitti e delle pene; Ferrajoli L. (1989), Diritto e ragione. Teoria del garantismo penale; Gospel of Luke 12:7; Court of Justice of the EU, Case C-131/12 (2014); Supreme Court of Wisconsin, State v. Loomis (2016); Constitution of the Italian Republic, Article 27; European Convention on Human Rights, Article 6; Directive (EU) 2016/343; Regulation (EU) 2016/679, Article 17; Regulation (EU) 2024/1689, Article 5, Articles 9 to 15, Annex III.

↑ Back to contents

by Antonio Fabbrizio · MMXXVI