IA on the job: assessing an output or a capability?

AI-assisted outputs are becoming part of performance reviews, promotion decisions and the allocation of new responsibilities. The management question is specific: what can we infer about a person’s capabilities from an output when its production conditions remain largely invisible?
What the literature establishes
Lee Ross, Teresa M. Amabile and Julia L. Steinmetz (1977, Journal of Personality and Social Psychology) studied a question-and-answer setting in which questioners could select questions drawing on their own knowledge. Participants did not sufficiently adjust their assessments of general knowledge to account for this role advantage. The experiment shows that an advantage created by a situation can be interpreted as a personal quality.
The mechanism: moving from output to person
Daniel T. Gilbert and Patrick S. Malone (1995, Psychological Bulletin) examine correspondence bias: the tendency to infer personal dispositions from behaviour that circumstances can explain. Their review distinguishes, among other mechanisms, a lack of awareness of constraints and insufficient correction of an initial judgement. Applied to AI, this framework suggests a hypothesis, not an established finding: managers may credit an employee for something that also depends on the tool, available resources and opportunities for revision.

The premature conclusion
One might conclude that AI should be removed from all assessments to reveal someone’s “real” capability. That would confuse individual proficiency without assistance with the ability to produce a reliable result in the actual working environment. Instead, the assessment target needs to be specified: the delivered output, subject-matter expertise or the ability to manage AI-assisted production. (our executive and employee training programmes)
What these studies cannot settle
These articles concern neither generative AI nor contemporary performance reviews, so they do not measure the size of any bias in those settings. A one-off demonstration without tools would likewise be insufficient to establish lasting competence. A limitation of current practice emerges when an assessment moves from a successful deliverable to a general judgement about its author without making that inference explicit.
A practical check in Berne
As part of SHR’s “IA on the job” programme, a team in Berne could test an assessment form using non-sensitive or fictional deliverables: a federal administration briefing, a public health document, a telecoms response or a deviation report in precision manufacturing. Before the exercise, the team would specify the capability being assessed; each assessor would then judge the deliverable alone and subsequently revise that judgement after receiving a standardised account of permitted tools, available resources and revisions made by its author. The measure would be the proportion of capability judgements changed after this contextual information, together with the reason for each change. This indicator would not prove bias, but would make it possible to check how much the assessment depends on information about production conditions. To go further: explore the AI on the Job training in Bern, or browse our executive and employee training programmes in Switzerland.
In pictures: AI on the Job in Bern



- ai on the job
- Bern
- research
- training bern
