NUS Computing and Stanford Researchers Co-Author Framework for Assessing AI’s Impact on Work
Researchers from NUS Computing, Stanford University and other institutions have co-authored a whitepaper setting out six conditions for judging whether AI genuinely improves work. The paper, “When Does AI Augment Work? A Workflow-Level Framework for Human-Agent Collaboration”, was posted to arXiv two days after the Ministry of Manpower told Parliament that its surveys show most firms adopting AI are redesigning jobs or creating new roles. Its co-authors are Jiaying Wu, a postdoctoral research fellow in Associate Professor Kan Min-Yen's WING.NUS research group at NUS Computing, and Caleb Ziems, a PhD student at Stanford University advised by Assistant Professor Diyi Yang.
The whitepaper grew out of CIVIC-AI 2026, a two-day workshop held at NUS on 13 and 14 July. Co-organised by faculty, staff and students from Associate Professor Kan's WING.NUS group and Asst Professor Yang’s SALT Lab at Stanford, the workshop was the first event in a new collaboration between the two universities on AI, work and society.
The first day was open to the public, with talks and a panel on social intelligence in AI systems. On the second day, an invitation-only session brought academics together with industry practitioners and regulators to discuss redesigning work with AI agents in Singapore. The paper draws on these second-day discussions.
The white paper notes that organisations currently track AI through usage logs, investment, productivity and headcount. The authors argue this tracking can miss whether a faster first draft has created extra checking or rework further along. The paper assesses whole workflows against six conditions: durable net value, meaningful human control, clear accountability and recovery, deepening learning, career pathways and job purpose.
The first three can be assessed in a snapshot evaluation but the last three conditions can only be assessed over time. The tasks most readily handed to AI tend to be high-volume, procedural and easy to verify, and junior workers build their judgement on the same tasks. A workflow can meet the first three conditions at launch and later depend on reviewers who never had that grounding.
The authors apply the framework to AI-assisted social surveys. In the division of labour proposed by workshop participants, an AI interviewer can rephrase questions, ask follow-ups and organise responses within limits the team has approved. Researchers define the study's purpose, decide the weighting, interpret the findings and approve the final coding. They also design the first protocols and do the initial coding themselves, partly because a researcher who starts from the AI's draft risks being anchored to it. That early coding is critical, as it is where junior researchers practise the fine-grained judgement they will later need for higher-level decisions.
“Staff need to be provisioned with the opportunity to do foundation work by hand, investing the necessary arduous investment in sensemaking. Without it, junior staff risk accumulating cognitive debt, as they get immediate feedback and critiques from fluent AI systems. This can lead to the collapse of having insufficient trained experienced staff, which itself further leads enterprises to rely on AI, which may have insufficient safeguards for unusual cases,” Prof Kan said.
The authors recommend keeping a shared record for each redesigned workflow. It would log what the agent does, where people review or override it, what goes wrong, and how the arrangement affects workers' skills and progression over time.
Read the full whitepaper here: https://arxiv.org/abs/2609.12482
