Skip to content
Research / perspective

What a human approval is worth

Research on automation says an approval step is weaker than it sounds. Why we keep ours, what our numbers show so far, and why we expect it to wear down.

Perspective6 min readiSyncSO · Team

The short version.

“A person approves” is one of the three rules SYNC works by. Under that rule a draft waits until someone has read it and released it.

The research on automation says that such a step is weaker than it sounds. People who work with a system that is usually right start to follow it. An approval can turn into a click.

We keep the rule, and we do not treat it as a guarantee. This piece sets out what the research says, what our own numbers show so far, and what we do to keep the approval a real decision.

What the research says about approval.

A review of the research on automation describes two errors that people make with a system that gives advice. They miss a problem because the system did not flag it, and they follow the system when it is wrong. The review is clear about who does this and whether it can be taught away:

The effect “occurs in both naive and expert participants, cannot be prevented by training or instructions”.
Parasuraman and Manzey, 2010, from the abstract[1]

Design can help, at a price. In an experiment with 199 people, designs that compel a person to engage with the advice before accepting it reduced over-reliance. The same people rated those designs lowest.[4] The design that works best is the one people like least.

The hardest criticism comes from a study of 41 policies that require a person to oversee a government algorithm. It finds that, on the evidence it reviews, people are unable to do the oversight asked of them, and that such policies can give “a false sense of security”.[5] In that reading the rule makes the system look safer than it is, because everyone assumes the person is checking.

Which part of the work the person keeps.

A classic paper divides any task into four kinds of function: gathering information, analysing it, deciding what to do, and doing it. Each can be automated to a different degree, from fully manual to fully automatic.[2] The useful question is which function is automated, and how far.

Gathering
In SYNCReading public sources about an account
Who does itThe software
Analysing
In SYNCScoring the account, building the case, mapping the people involved
Who does itThe software
Deciding
In SYNCThe software proposes one draft. Whether it is sent is decided at release
Who does itA person, unless approval was relaxed for the account
Doing
In SYNCSending the released message from your own mailbox
Who does itThe software

The division has a known weakness. An older paper on automation in industry argues that when most of a task is automated, the person is left with watching and with the hard cases, and has less practice for both. We know this paper from summaries and have not read it ourselves.[3]

That describes our reviewer. They did not gather the sources and did not write the draft. They are asked to judge a text they had no hand in, and to catch the one draft in many that is wrong.

What our own numbers show.

We compared the first draft with the sent text for every draft a person released between 21 August and 29 September 2026.

Measured by us · 29 September 2026 · n = 266 released drafts
249of 266 released drafts left with a changed text (94%)
17of 266 were released exactly as written

Most of the changed drafts got longer, and most were rewritten to a large extent. The full count is in What reps change before they send.[8]

So far our reviewers are not approving on sight. The limits are large. The reviewers were a small group, and the reviews are spread unevenly across them. Some of the change may be the software's own repair, which we cannot separate. And a change says nothing about whether the facts in the message were checked.

What we do about it.

  • 01The reviewer sees the reason for writing next to the draft: the fact the message rests on, a link to its source, and the date it was observed. When the fact has no date, the screen says so. The reviewer can check the claim, and not only the tone.
  • 02The reviewer also sees which statements about their own company the draft relied on that nobody in the workspace has confirmed yet.
  • 03Approval is set per account, at one of three levels: every message, the first message only, or none. Today a new account starts with approval of the first message. Follow-ups can then go out without review, unless the workspace chose approval on every send. We are changing this, so that every new account starts with approval on every send. [Date of the change to publish.]
  • 04Since 29 September the software stores what changed between the draft and the sent text, so that we can follow over time whether reviewers keep editing.
  • 05Today the review screen shows the reason for writing, its source, and its date or that the date is missing. Showing the source under each sentence of the draft is being built. We give no date for it.

One finding belongs here, because it touches the first item. When we measured, this site stated the rule that an account without a dated reason gets nothing from SYNC. On 29 September, 570 dossiers were marked ready. By the reason recorded with that verdict, 303 rest on a dated, cited signal, 230 on a cited signal without a date, and 37 have no reason recorded.[8] Today the check asks for a cited signal and its source, and does not require a date. We decided to keep this: a cited signal without a date can still make a dossier ready. The reviewer sees that the date is missing, and we are making that more visible in the review screen. The wording of the rule on this site now says a reason we can cite, not a dated reason.

The European AI Act gives a useful benchmark for the design. Its article on human oversight asks that a person who oversees a system stays aware of “the possible tendency of automatically relying or over-relying on the output”.[7] That article is written for systems the Act calls high-risk. Sales outreach is not on the list of high-risk uses, so the article very likely does not apply to SYNC as law. We use it as a standard to design against. We make no claim that SYNC complies with the Act, or that it needs to.

the other side

What speaks against this.

The research we cite predicts that our own approval step will wear down. Reviewers who see good drafts for months will start to trust them. The high share of edits we measure today may fall, and a falling share could mean better drafts or less attention. We do not yet measure which, and we do not yet follow the share over time.

Showing the source does not make anyone open it. The authors of one study note that adding explanations to advice does not appear to reduce over-reliance.[4] A source is not an explanation, but the warning carries over. We do not record whether a reviewer followed the link to a source before releasing.

Letting the person relax approval per account is the opposite of a design that forces thought. The research says people prefer the designs that protect them least, so we expect the relaxed settings to be used.[4] An account that runs without approval has no human check at all.

The research is not of one mind, and none of it is about sales. A review of studies in health care lists training and stressing the user's own accountability among the things that reduce the effect.[6] That sits badly with the claim that training cannot prevent it. We read most of these papers in the abstract only, and one through summaries.

Our own evidence is a small group of reviewers over less than six weeks. It cannot carry a general claim about what an approval is worth.

method · measured 29 September 2026

The count of released drafts comes from the log of every draft and what happened to it: 266 drafts that a person reviewed and released between 21 August and 29 September 2026, by a small group of reviewers. 24 drafts from a second writing route were left out, because their stored first draft is a different draft. The count of dossiers, the research files that SYNC keeps per account, comes from 1,303 dossiers in five workspaces, of which 570 were marked ready. For those 570 we read the reason that the software recorded with the verdict.

We report all workspaces together and never one workspace on its own. The figures are valid for the moment they were taken. We measure again on the day of publication.

How we measure →
sources
  1. [1]Parasuraman, R. & Manzey, D. H. (2010). Complacency and bias in human use of automation: an attentional integration. Human Factors 52(3), 381–410.abstract only
  2. [2]Parasuraman, R., Sheridan, T. B. & Wickens, C. D. (2000). A model for types and levels of human interaction with automation. IEEE Transactions on Systems, Man, and Cybernetics, Part A 30(3), 286–297.abstract only
  3. [3]Bainbridge, L. (1983). Ironies of automation. Automatica 19(6), 775–779.secondaryknown to us from summaries, not from the paper itself
  4. [4]Buçinca, Z., Malaya, M. B. & Gajos, K. Z. (2021). To trust or to think: cognitive forcing functions can reduce overreliance on AI in AI-assisted decision-making. Proceedings of the ACM on Human-Computer Interaction (CSCW).abstract onlyread on arXiv; an experiment with 199 people, not about sales
  5. [5]Green, B. (2022). The flaws of policies requiring human oversight of government algorithms. Computer Law & Security Review 45.abstract onlyabout government algorithms
  6. [6]Goddard, K., Roudsari, A. & Wyatt, J. C. (2012). Automation bias: a systematic review of frequency, effect mediators, and mitigators. Journal of the American Medical Informatics Association 19(1), 121–127.abstract onlyabout decision support in health care
  7. [7]Regulation (EU) 2024/1689 (Artificial Intelligence Act), article 14 and Annex III.read in the originalused as a design benchmark; we make no compliance claim
  8. [8]iSyncSO. Drafts released after a person's review, and the verdicts of researched accounts. Measured 29 September 2026.our own measurementsee the method note

See who should buy from you this quarter.

Free. We open in January.