Skip to content
Research / research

Advice a model may skip is advice it does skip

In our clearest example, a rule that was only stated to the writing model drifted. The rules that matter are now checks that refuse a draft. What those checks caught, and what that does not prove.

Research4 min readiSyncSO · Team

The short version.

There are two ways to give a rule to the model that writes a draft. You can state the rule in the instructions and trust the model to follow it. Or the software can check the finished draft and refuse it when the rule is broken.

In our clearest example the stated rule drifted. For about two weeks the instructions told the model exactly how to sign off. Of 178 drafts in that period, 40 had the sign-off exactly as instructed, 93 differed from it and 45 had none.

So the rules that matter are now checks that refuse, and every refusal is logged with the rule that caused it. This piece shows what those checks caught. It does not show that the messages became better, and the example is one configuration, not a general rate.

One instruction, unchecked for two weeks.

A workspace can set how its messages end: the closing words and the name. It is the simplest rule we have. It needs no judgement, only copying.

From about 7 to 20 September the instructions to the model named the sign-off, and nothing checked the result. We looked at the model's own drafts from that period and compared the ending with the sign-off that was set.

Measured by us · 29 September 2026 · n = 178 drafts
40of 178 drafts had the sign-off exactly as set (22%)
93of 178 had a sign-off that differed from the one set (52%)
45of 178 had no sign-off at all (25%)

Three limits belong with these numbers. Every draft counted as “differs” comes from one configuration and differs in the same small way. We compared old drafts with the sign-off as it is set today. And the start of the period is the date the instruction was written, not the date it went live. So this shows one configuration drifting. It is not a general rate for how often a model skips an instruction.[1]

On the evening of 20 September the sign-off became a check. A draft without the right sign-off is refused. Since then the check has refused 46 times: 44 times because the sign-off was missing and twice because it differed.

These 46 refusals cannot be set against the 178 drafts. A refusal is an event, and one draft can be refused several times. The two counts measure different things in different periods.

What the checks caught.

SYNC writes along two routes. In the first, 527 refusals were logged between 7 and 29 September. The table sorts them by what the refusal was about.

Measured by us · 29 September 2026 · n = 527 refusals
The text has no paragraph breaks
Refusals86
First seen11 Sept
The message names its source, or says that someone looked
Refusals65
First seen7 Sept
A number about the sender that the sender's own description does not contain
Refusals63
First seen9 Sept
The closing question asks for a time or a form the writing rules do not allow
Refusals60
First seen7 Sept
The sign-off is missing
Refusals44
First seen20 Sept
A follow-up repeats the earlier message below its opening
Refusals42
First seen9 Sept
Sentences repeat an email that was already sent
Refusals39
First seen7 Sept
The draft came back with a part missing
Refusals25
First seen11 Sept
A date that the research on the account does not support
Refusals24
First seen11 Sept
A first word that the writing rules forbid
Refusals22
First seen9 Sept
The answer could not be read as a draft
Refusals20
First seen7 Sept
The same argument as in recent messages to other accounts
Refusals11
First seen20 Sept
The message lectures the reader
Refusals8
First seen9 Sept
A phrase that the writing rules forbid
Refusals7
First seen21 Sept
The message switches between formal and informal address
Refusals5
First seen25 Sept
The sign-off differs from the one set
Refusals2
First seen23 Sept
Other
Refusals4
First seenvarious

This is not a ranking of the model's mistakes. The checks were added week by week, as the last column shows. A check that exists for three weeks has had more time to refuse than one that exists for four days. A mistake for which no check exists does not appear at all.

The second route records every attempt to write, from 21 August on. It shows where the effort goes.

Measured by us · 29 September 2026 · n = 221 writing attempts
Attempts to write a draft
Count221
Accepted drafts
Count54
Attempts that failed a check
Count133
Of those, stopped first by the check on form
Count122
Of those, stopped first by a style check
Count11
Attempts where the connection failed
Count34

By our own division, that is about four attempts for every accepted draft. Most failures have nothing to do with content. In 122 of the 133 failed checks the model did not return its answer in the form that was asked for, so there was no draft to judge.

What outside research says.

We found no study of this question for sales email. The nearest research asks whether a model that is told to support its statements with sources does so. Two studies measured it.

“even the best models lack complete citation support 50% of the time”
Gao, Yen, Yu and Chen, 2023[2]

That result is for one set of long-form questions. A second study checked four public search engines that answer in full sentences with sources. It found that on average “a mere 51.5% of generated sentences are fully supported by citations”.[3]

Giving a model the source material does help. One study found that models which look up passages before they write produce more factual language than models which do not.[4] Another found that looking up passages substantially reduces invented facts in conversation.[5] Both report a reduction. Neither reports that the problem is gone.

What we changed in the product.

  • 01A rule that matters is a check on the finished draft. When the draft breaks the rule, the software refuses it.
  • 02Every refusal is logged with the rule that caused it. That log is the source of the tables in this piece.
  • 03The sign-off moved from the instructions to a check on 20 September.
the other side

What speaks against this.

We cannot show that the checks improved the messages. We count refusals, and a refusal only says that a pattern matched. A refused draft may have been fine in a person's eyes, and a draft that passed may still be poor.

A check can only test what software can test. Paragraph breaks, a sign-off and a forbidden word are easy. Whether a message is relevant to its reader is not. The most important qualities of a message stay advice, and by our own finding advice drifts.

Checks cost attempts. In the second route, 44 of 98 runs never produced an accepted draft. Each refused attempt costs time and money, and a strict set of checks can leave an account without a message.

Our clearest example is narrow. The sign-off numbers come from very few configurations, and all the differing drafts from one of them. The start of the period is approximate: it is the date on which the instruction was added, not a date recorded with each draft. All our counts come from three workspaces and a few weeks.

The outside studies are about answering questions and about conversation. None of them studied sales email, and none compared a stated rule with a checked rule. We use them as the nearest evidence, not as proof.

method · measured 29 September 2026

We counted refusals in the log of every draft and what happened to it: 527 in the first writing route, from 7 to 29 September, and 133 failed checks in the second, from 21 August to 29 September. We report all workspaces together. We give no refusal rate, because refusals are events and one draft can be refused more than once.

The categories in the first table are made by matching patterns in the text of the reason that was logged with each refusal. For the sign-off we looked at the model's own drafts in the first writing route, in workspaces that have set a sign-off. Left out: drafts from the second writing route, drafts written after the check began, and drafts from before 7 September, because the start of the period is approximate and we cannot say which instructions those drafts were written under.

The figures are valid for the moment they were taken. We measure again on the day of publication.

How we measure →
sources
  1. [1]iSyncSO. Decisions on drafts, refusals by rule and writing attempts, to 29 September 2026. Measured 29 September 2026.our own measurementa small number of workspaces
  2. [2]Gao, T., Yen, H., Yu, J. & Chen, D. (2023). Enabling large language models to generate text with citations. arXiv:2305.14627.abstract onlyread on arXiv; question answering, not sales email
  3. [3]Liu, N. F., Zhang, T. & Liang, P. (2023). Evaluating verifiability in generative search engines. arXiv:2304.09848.abstract onlyread on arXiv; four search engines, not sales email
  4. [4]Lewis, P. et al. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. NeurIPS 2020. arXiv:2005.11401.abstract onlyread on arXiv
  5. [5]Shuster, K., Poff, S., Chen, M., Kiela, D. & Weston, J. (2021). Retrieval augmentation reduces hallucination in conversation. arXiv:2104.07567.abstract onlyread on arXiv

See who should buy from you this quarter.

Free. We open in January.