The short version.
SYNC writes a draft, and a person decides whether it is sent. We compared what the software first wrote with what was finally sent. Of 266 drafts that a person released, 249 left with a different text. 17 went out as written.
Most changes are large, and most make the message longer. That does not look like approval on sight.
The limits are large. The numbers describe a small group of reviewers, and the reviews are spread unevenly across them. Part of the change may be the software's own repair of a draft. And nothing here shows that the changes made a message better, because we measure no outcome.
How many drafts were changed.
Every one of the 249 had a change in the body of the message. No change was only a matter of spaces or line breaks.
The drafts were released between 21 August and 29 September 2026. A small group of people did the reviewing.[1]
We started with 290 released drafts and left out 24. Those 24 came from a second writing route in which the stored first draft is a different draft from the one that was reviewed, so comparing the two would measure nothing.
How large the changes are.
The changes are not small corrections. We measured them in two ways: how much the length changed, and how similar the sent text still is to the draft.
| What we measured | Result |
|---|---|
| Got longer | 214 of 249 |
| Got shorter | 34 of 249 |
| Change in length, middle value | 24% longer |
| Change in length, middle half of the drafts | between 8% and 55% longer |
Similarity is a score from 0 to 1. It compares the two texts in overlapping groups of three characters. A score of 1 means the texts are the same, and a score near 0 means they share almost nothing.
Adding the two lowest groups, 189 of the 249 changed drafts score below 0.70. That is three in four. Only 18 were touched lightly. The sum and the descriptions next to the scores are ours. The score itself only says how much of the text is shared.
We cannot yet say what kind of changes these are. Since 29 September the software records a comparison at each release and sorts the differences into four kinds: facts, who the message is for, structure and tone, and the closing. There are eleven such records, all from a single day and a very small group. Each of the eleven carries all four kinds, so the sorting separates nothing yet.
What this suggests, and what it does not.
The research on automation gives a reason to look at this number. People who work with a system that is usually right tend to follow it. A review of that research says the effect appears in beginners and in experts, and that it “cannot be prevented by training or instructions”.[2] An approval step where almost everything passes unchanged would fit that picture.
Ours does not look like that today. 17 of 266 drafts passed unchanged. The reviewers change the text, and they add to it more often than they cut.
- 01It does not show that the changes improved the message. We measure no replies or meetings in this piece, and we do not publish such rates.
- 02It does not show that the draft was wrong. A person can rewrite a correct draft into their own voice.
- 03It does not show that an unchanged draft was approved without reading. A person can read a draft carefully and find nothing to change.
- 04It does not describe sales people in general. It describes a small group of reviewers.
What we changed in the product.
We added the comparison at release that is described above. Since 29 September the software stores, at release, what changed between the draft and the sent text. Once enough records exist, we can report which kinds of change are common.
We have not yet changed what is stored before the review. The limits below follow from that.
What speaks against this.
Some of the change may be the software's own. After the first draft, the software can repair a draft that breaks one of its rules, and only then does a person see it. We store the first draft and the sent text. We do not store the text the reviewer read. For a released draft we cannot separate the repair from the person's edit, so the person's share may be smaller than 249 of 266, by an amount we do not know.
The numbers are small and uneven. A small group of people reviewed, over less than six weeks, and the reviews are spread unevenly across them. The result describes how a few people work.
A high share of edits can also be read as a weakness of the drafts. If 94% of drafts need rewriting, the software is saving less work than it seems to. Our data cannot decide between that reading and ours.
Editing is also not the same as checking. A study of 41 policies that require a person to oversee a government algorithm finds that, on the evidence it reviews, people are unable to do the oversight asked of them, and that such policies can give “a false sense of security”.[4] A reviewer can improve the wording of a message and still miss a wrong fact in it. One experiment found that designs which force a person to think reduce over-reliance, and that people rate those designs lowest.[3] Both studies are about other fields than sales.
The share may also fall. On 19 September the same count stood at 113 of 126. It moves with every release, and people who come to trust the drafts may edit less. We do not yet follow that over time.
We counted drafts that a person reviewed and released, in the log of every draft and what happened to it. The period runs from 21 August to 29 September 2026 and covers a small group of reviewers. We compared the body and the subject of the first draft with the body and the subject of the sent message.
Left out: 24 drafts from the second writing route, because their stored first draft is a different draft. Drafts still waiting for review, because their changes so far are the software's own repairs. We report all workspaces together and never one workspace on its own. 189 and the rounded shares are our own sums from the counts above.
The figures are valid for the moment they were taken. We measure again on the day of publication.
- [1]iSyncSO. Drafts released after a person's review, 21 August to 29 September 2026. Measured 29 September 2026.our own measurementa small group of reviewers
- [2]Parasuraman, R. & Manzey, D. H. (2010). Complacency and bias in human use of automation: an attentional integration. Human Factors 52(3), 381–410.abstract only
- [3]Buçinca, Z., Malaya, M. B. & Gajos, K. Z. (2021). To trust or to think: cognitive forcing functions can reduce overreliance on AI in AI-assisted decision-making. Proceedings of the ACM on Human-Computer Interaction (CSCW).abstract onlyread on arXiv; an experiment with 199 people, not about sales
- [4]Green, B. (2022). The flaws of policies requiring human oversight of government algorithms. Computer Law & Security Review 45.abstract onlyabout government algorithms