Lesson 1 of 7 · 7 min

Where bias comes from

The biases of an AI system come from the data, from design choices and from how it is used. Two studies have shown this clearly.

ObjectiveIdentify the three sources of bias and illustrate them with documented cases.

An AI system learns from examples. If the examples are imbalanced, if they reflect past inequalities, or if the chosen objective poorly measures what you actually want, the system reproduces these flaws at scale, with an appearance of neutrality.

First case: imbalanced data

In 2018, the Gender Shades study evaluated three commercial systems that classify gender from photos. The error rate reached 34.7% for darker-skinned women, against at most 0.8% for lighter-skinned men. The datasets used consisted overwhelmingly of lighter-skinned people.

Second case: a poorly chosen objective

In 2019, a study published in Science analysed an algorithm used in the United States to refer patients to enhanced care programmes. It estimated care needs from past healthcare spending. Yet, for the same state of health, less money had been spent on Black patients. At the same score, they were therefore sicker. Correcting the objective raised their share among referred patients from 17.7% to 46.5%.

The three sources to check

  • The data: who is represented, who is missing, and how old the data is.
  • The objective: does the predicted quantity really measure what matters?
  • The use: is the system being applied to a population or a decision it was not designed for?

References

  1. Buolamwini, Gebru (2018). Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification. Conference on Fairness, Accountability and Transparency, PMLR 81, p. 77–91. proceedings.mlr.press/v81/buolamwini18a.html
  2. Obermeyer, Powers, Vogeli et al. (2019). Dissecting racial bias in an algorithm used to manage the health of populations. Science 366(6464), p. 447–453. doi.org/10.1126/science.aax2342