The short version.
There are two ways to give a rule to the model that writes a draft. You can state the rule in the instructions and trust the model to follow it. Or the software can check the finished draft and refuse it when the rule is broken.
In our clearest example the stated rule drifted. For about two weeks the instructions told the model exactly how to sign off. Of 178 drafts in that period, 40 had the sign-off exactly as instructed, 93 differed from it and 45 had none.
So the rules that matter are now checks that refuse, and every refusal is logged with the rule that caused it. This piece shows what those checks caught. It does not show that the messages became better, and the example is one configuration, not a general rate.
One instruction, unchecked for two weeks.
A workspace can set how its messages end: the closing words and the name. It is the simplest rule we have. It needs no judgement, only copying.
From about 7 to 20 September the instructions to the model named the sign-off, and nothing checked the result. We looked at the model's own drafts from that period and compared the ending with the sign-off that was set.
Three limits belong with these numbers. Every draft counted as “differs” comes from one configuration and differs in the same small way. We compared old drafts with the sign-off as it is set today. And the start of the period is the date the instruction was written, not the date it went live. So this shows one configuration drifting. It is not a general rate for how often a model skips an instruction.[1]
On the evening of 20 September the sign-off became a check. A draft without the right sign-off is refused. Since then the check has refused 46 times: 44 times because the sign-off was missing and twice because it differed.
These 46 refusals cannot be set against the 178 drafts. A refusal is an event, and one draft can be refused several times. The two counts measure different things in different periods.
What the checks caught.
SYNC writes along two routes. In the first, 527 refusals were logged between 7 and 29 September. The table sorts them by what the refusal was about.
| What the check refused | Refusals | First seen |
|---|---|---|
| The text has no paragraph breaks | 86 | 11 Sept |
| The message names its source, or says that someone looked | 65 | 7 Sept |
| A number about the sender that the sender's own description does not contain | 63 | 9 Sept |
| The closing question asks for a time or a form the writing rules do not allow | 60 | 7 Sept |
| The sign-off is missing | 44 | 20 Sept |
| A follow-up repeats the earlier message below its opening | 42 | 9 Sept |
| Sentences repeat an email that was already sent | 39 | 7 Sept |
| The draft came back with a part missing | 25 | 11 Sept |
| A date that the research on the account does not support | 24 | 11 Sept |
| A first word that the writing rules forbid | 22 | 9 Sept |
| The answer could not be read as a draft | 20 | 7 Sept |
| The same argument as in recent messages to other accounts | 11 | 20 Sept |
| The message lectures the reader | 8 | 9 Sept |
| A phrase that the writing rules forbid | 7 | 21 Sept |
| The message switches between formal and informal address | 5 | 25 Sept |
| The sign-off differs from the one set | 2 | 23 Sept |
| Other | 4 | various |
This is not a ranking of the model's mistakes. The checks were added week by week, as the last column shows. A check that exists for three weeks has had more time to refuse than one that exists for four days. A mistake for which no check exists does not appear at all.
The second route records every attempt to write, from 21 August on. It shows where the effort goes.
| Second writing route | Count |
|---|---|
| Attempts to write a draft | 221 |
| Accepted drafts | 54 |
| Attempts that failed a check | 133 |
| Of those, stopped first by the check on form | 122 |
| Of those, stopped first by a style check | 11 |
| Attempts where the connection failed | 34 |
By our own division, that is about four attempts for every accepted draft. Most failures have nothing to do with content. In 122 of the 133 failed checks the model did not return its answer in the form that was asked for, so there was no draft to judge.
What outside research says.
We found no study of this question for sales email. The nearest research asks whether a model that is told to support its statements with sources does so. Two studies measured it.
“even the best models lack complete citation support 50% of the time”
That result is for one set of long-form questions. A second study checked four public search engines that answer in full sentences with sources. It found that on average “a mere 51.5% of generated sentences are fully supported by citations”.[3]
Giving a model the source material does help. One study found that models which look up passages before they write produce more factual language than models which do not.[4] Another found that looking up passages substantially reduces invented facts in conversation.[5] Both report a reduction. Neither reports that the problem is gone.
What we changed in the product.
- 01A rule that matters is a check on the finished draft. When the draft breaks the rule, the software refuses it.
- 02Every refusal is logged with the rule that caused it. That log is the source of the tables in this piece.
- 03The sign-off moved from the instructions to a check on 20 September.
What speaks against this.
We cannot show that the checks improved the messages. We count refusals, and a refusal only says that a pattern matched. A refused draft may have been fine in a person's eyes, and a draft that passed may still be poor.
A check can only test what software can test. Paragraph breaks, a sign-off and a forbidden word are easy. Whether a message is relevant to its reader is not. The most important qualities of a message stay advice, and by our own finding advice drifts.
Checks cost attempts. In the second route, 44 of 98 runs never produced an accepted draft. Each refused attempt costs time and money, and a strict set of checks can leave an account without a message.
Our clearest example is narrow. The sign-off numbers come from very few configurations, and all the differing drafts from one of them. The start of the period is approximate: it is the date on which the instruction was added, not a date recorded with each draft. All our counts come from three workspaces and a few weeks.
The outside studies are about answering questions and about conversation. None of them studied sales email, and none compared a stated rule with a checked rule. We use them as the nearest evidence, not as proof.
We counted refusals in the log of every draft and what happened to it: 527 in the first writing route, from 7 to 29 September, and 133 failed checks in the second, from 21 August to 29 September. We report all workspaces together. We give no refusal rate, because refusals are events and one draft can be refused more than once.
The categories in the first table are made by matching patterns in the text of the reason that was logged with each refusal. For the sign-off we looked at the model's own drafts in the first writing route, in workspaces that have set a sign-off. Left out: drafts from the second writing route, drafts written after the check began, and drafts from before 7 September, because the start of the period is approximate and we cannot say which instructions those drafts were written under.
The figures are valid for the moment they were taken. We measure again on the day of publication.
- [1]iSyncSO. Decisions on drafts, refusals by rule and writing attempts, to 29 September 2026. Measured 29 September 2026.our own measurementa small number of workspaces
- [2]Gao, T., Yen, H., Yu, J. & Chen, D. (2023). Enabling large language models to generate text with citations. arXiv:2305.14627.abstract onlyread on arXiv; question answering, not sales email
- [3]Liu, N. F., Zhang, T. & Liang, P. (2023). Evaluating verifiability in generative search engines. arXiv:2304.09848.abstract onlyread on arXiv; four search engines, not sales email
- [4]Lewis, P. et al. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. NeurIPS 2020. arXiv:2005.11401.abstract onlyread on arXiv
- [5]Shuster, K., Poff, S., Chen, M., Kiela, D. & Weston, J. (2021). Retrieval augmentation reduces hallucination in conversation. arXiv:2104.07567.abstract onlyread on arXiv