Errors in Errors: An Exploration of Grammarly’s Corrective Feedback
About this paper:
- Joshua Kloppers published this research article in volume 13, issue 1 of the International Journal of Computer-Assisted Language Learning and Teaching in 2023.
- Using ratings and interviews from two native English speakers, the mixed-methods study evaluates the accuracy and quality of Grammarly feedback on English texts written by Chinese-speaking undergraduates.
- Its central implication is that Grammarly can assist with formal errors but its style suggestions should not be treated as authoritative corrections, making learner training and critical selection essential.
Introduction
Automated Writing Evaluation (AWE) tools such as Grammarly can flag an error and propose a correction even when no teacher is available. This immediacy is especially attractive to second-language writers. Students can receive feedback before submitting a text, while teachers can spend less time correcting recurring formal errors. Fast feedback, however, is not necessarily accurate feedback.
The paper does not ask whether using Grammarly improves writing ability. Instead, it treats each error flag and attached correction as an attentional cue and evaluates that cue on its own terms. Rather than comparing an entire Grammarly report with an instructor’s complete marking, it examines what learners actually encounter on screen: Is the highlighted passage really erroneous, and is the suggested replacement appropriate?
This distinction determines how the results should be read. Studying only items that Grammarly flags can measure the quality of those flags, including false positives. It cannot reveal errors in the source texts that Grammarly never detected. The paper’s “accuracy” is therefore closer to the precision or trustworthiness of displayed feedback than to total error-detection performance. A high percentage does not imply high recall.
The author asks two questions:
- According to native English speakers, how accurate are Grammarly’s error identifications and corrections in L2 writing?
- How do those evaluators perceive and describe the identified errors and suggested corrections?
The first question is addressed quantitatively through coded ratings, and the second qualitatively through interviews and examples. The work is thus a mixed-methods evaluation.
The product’s date must also remain visible. The article was published in 2023 and analyzes a report generated by the then-current Grammarly Premium. Commercial AWE systems change frequently, so these percentages cannot simply be assigned to the 2026 product. The more durable contribution is the paper’s demonstration of how automated feedback should be evaluated, not a permanent score for Grammarly.
Main Discussion
The Grammarly report, not the student essays, is the analytical dataset
The author randomly selected forty texts with clear evidence of undergraduate authorship from the Ten-thousand English Compositions of Chinese Learners Corpus (TECCL). TECCL contains classroom and test writing by English learners in mainland China, Hong Kong, and Taiwan. The selected texts contained 6,286 words.
Two experienced English teachers placed the texts at CEFR A2, while the automated CatHoven analyzer rated them between B1 and B2. The author uses this disagreement as another example of human and automated text assessment producing different judgments. The level estimate is not a main analytical variable, but it establishes that the source material represents relatively low-proficiency undergraduate L2 writing.
The forty texts were merged into one document and processed with Grammarly Premium. The settings were “knowledgeable” audience, “neutral” formality, and “general” domain. The report contained 511 sentences, a mean word length of 4.4 characters, a mean sentence length of 12.3 words, and a Grammarly score of 25/100.
The abstract, method, and conclusion repeatedly state that the process generated 1,136 Grammarly-identified errors. The Results section and Tables 3–4, however, consistently total 1,126. The 257 style items and 869 correctness items also sum to 1,126. The paper does not explain the ten-item discrepancy. Because all reported percentages reproduce correctly with 1,126 as their denominator, the result tables should be interpreted using 1,126, while the inconsistency itself should be recorded as a reporting error.
Identification and correction were evaluated separately
The two raters were native English speakers from South Africa. Both were university graduates with substantial academic writing experience: one had a medical background and the other a psychology background, while only the latter had three years of English-language teaching experience.
The coding scheme separates Grammarly’s error identification from its correction. An identification received one of four judgments:
- an error was accurately identified and required correction;
- the highlighted text was not erroneous;
- a problem existed, but Grammarly identified its nature or extent misleadingly;
- the rater could not determine the status.
The correction received one of five judgments:
- appropriate and necessary;
- possible but unnecessary, incomplete, or suboptimal;
- inappropriate and harmful to the sentence;
- no correction supplied;
- impossible for the rater to judge.
A code such as 1A means that both the identification and proposed correction were appropriate. 1D means that an error was correctly identified but no correction was supplied. The paper counts both as accurate in its overall measure. In other words, a correct flag can qualify as accurate feedback even without a proposed repair.
After jointly practicing on a separate training text with the researcher, the raters independently coded the 1,126 items. Their pre-consensus disagreement rate was 6.67%, producing Krippendorff’s α = 0.72 with a 95% confidence interval of [0.65, 0.78]. This is below the frequently used 0.80 threshold. Each rater then completed a 45-minute semi-structured interview. They met to resolve every disagreement and participated in a final 15-minute joint interview.
Consensus is practical for producing a final table, but it also matters analytically. The final counts are negotiated judgments, not the raw agreement of two independent raters. Style is especially subjective, so consensus between two people cannot define a universal standard of good English prose.
The reported 78.86% changes when its components are separated
Table 3 yields the following high-level distribution.
| Rating outcome | Items | Share of all items |
|---|---|---|
| Accurate identification + appropriate correction | 808 | 71.76% |
| Accurate identification + no correction | 80 | 7.10% |
| All other combinations | 238 | 21.14% |
| Total | 1,126 | 100% |
The paper combines the first two rows, obtaining $888/1,126=78.86\%$. Different questions produce different values:
- If only correct identification is considered, the value is 931/1,126, or 82.68%.
- Among the 958 items that include a correction, 808 have both an accurate identification and an appropriate correction, or 84.34%.
- Among the 168 items without a correction, only 80 are accurately identified, which gives a table-derived accuracy of 47.62%.
The third number has immediate pedagogical relevance. When Grammarly flagged a problem but could not suggest a replacement, more than half of those flags were misleading or inaccurate according to the consensus ratings.
The Discussion contains a confusing sentence that presents 78.86% and 85.07% as though they were the accuracies of cues with and without corrections. In fact, 85.07% is the proportion of all cues that had corrections, $958/1,126$, not an accuracy rate. Reading Table 3 directly gives 84.34% fully accurate cues among items with corrections and 47.62% accurate identifications among items without corrections.
Correctness scored 91.60%, while style scored 35.80%
The difference between feedback domains is more informative than the overall average.
| Grammarly domain | Items | Rated accurate | Accuracy |
|---|---|---|---|
| Style | 257 | 92 | 35.80% |
| Correctness | 869 | 796 | 91.60% |
Correctness includes comparatively rule-governed problems involving grammar, spelling, punctuation, and formatting. No correctness subtype with more than ten instances fell below 70%. Notable results included improper formatting at 99.36%, incorrect noun number at 95.65%, wrong or missing prepositions at 94.29%, conjunction use at 94.12%, and faulty tense sequence at 91.30%. Punctuation in compound or complex sentences was weaker at 76.54%, but still substantially stronger than the style category overall.
Style comprises clarity, engagement, and delivery, whose accuracies were 52.94%, 23.81%, and 16.67%, respectively. At the finer level, wordy sentences scored 71.19% and monotonous sentences 85.71%, while passive voice misuse scored 23.53%, word choice 18.18%, and tone suggestions only 12.00%. Percentages based on one or two cases, including some apparent 100% values, should not be generalized without their sample sizes.
This gap disappears if the result is compressed to “Grammarly is 78.86% accurate.” A user should assign very different prior trust to a spelling or noun-number flag than to a recommendation about tone, wording, passive voice, or concision.
The examples reveal failures of context and meaning preservation
The interviews show how low style accuracy appears in practice.
- Grammarly proposed changing
rich heritagetowealthy heritageto make the prose more engaging, although the expressions do not preserve the same meaning. - It tried to change
Pan, the name of a character in a novel, to lowercasepanto make it consistent with references to the kitchen utensil. - It replaced
satisfiedwithdelightedfor lexical variety, strengthening the sentiment beyond what the writer may have intended. - It labeled
We are prepared to give ...as passive voice misuse, even though the raters considered the construction natural. - For
In my opinion. I think it's better..., it left the erroneous period in place and proposedIn, my opinion., inserting a comma after “In.”
Each failure requires context or authorial intent beyond a narrow local pattern. The raters judged Grammarly stronger at phrase- and clause-level correction and weaker as the required context expanded to the sentence or discourse level.
Both raters independently felt that Grammarly promoted a traditional, formal, academic style despite the use of the “general” domain setting. They also worried that suggestions intended to increase engagement could homogenize writing into a predictable formula and make it less engaging. Style feedback can therefore do more than produce a false positive: it can alter voice, stance, and the learner’s sense of ownership.
Automated feedback should be treated as a prompt to review, not a command to revise
The paper does not recommend abandoning Grammarly. Correctness feedback was generally accurate, and immediate assistance may be valuable when instructor feedback is unavailable. Blind acceptance, however, can teach students to avoid any form that receives a warning rather than understand why a revision is or is not appropriate.
One rater suggested using style flags as attentional devices: not “this is wrong, so replace it,” but “reread this passage and decide whether it achieves your purpose.” This interpretation preserves a possible educational benefit without pretending that the style judgment is reliable. It also requires training learners to articulate why they accept or reject a suggestion.
I would translate the findings into five practical rules:
- Inspect formal, spelling, and clear grammatical flags first, while still comparing them with the original.
- Assign especially low confidence to warnings without a concrete correction.
- Convert vocabulary, tone, passive voice, and concision suggestions into review questions rather than answers.
- Check whether the revision preserves semantic strength, authorial intent, and genre.
- Require the learner to justify acceptance with grammar or context instead of appealing to the tool’s authority.
Scope and limitations
The paper’s stated limitations and additional constraints visible in the data include the following.
- There were only two raters, and their initial α of 0.72 did not demonstrate high reliability.
- Only one rater had language-teaching experience, and both were native English speakers, limiting representation of L2 intentions and plural English norms.
- The dataset contained only forty texts and 6,286 words from Chinese-speaking undergraduates; other proficiency levels, first languages, and genres may produce different results.
- Merging all essays into one document may have affected how document boundaries and context were processed.
- Because undetected source-text errors were never annotated, recall and total detection performance cannot be calculated.
- Subcategory frequencies vary sharply, and several percentages are based on only one to seven cases.
- The paper alternates between 1,136 and 1,126 items and contains a confusing description of the Table 3 percentages.
- The study captures one configuration of a rapidly updated commercial product and cannot establish the performance of its current version or other settings.
These constraints mean that the paper does not justify replacing instructors with Grammarly. They also do not show that all automated feedback is worthless. The best-supported claim is narrower and more useful: trust should vary by feedback category.
Conclusion
Before reading the study, it might seem reasonable to evaluate Grammarly with one accuracy number. Once the result is decomposed, however, the underlying questions separate. Did the system highlight the right span? Did it describe the problem correctly? Is the proposed sentence grammatical? Does the revision preserve the writer’s meaning and fit the genre? The overall 78.86% conceals these distinctions.
The clearest finding is the contrast between 91.60% correctness accuracy and 35.80% style accuracy. Grammarly was useful for local, comparatively rule-governed formal errors but unreliable when feedback required semantic, discourse, or authorial judgment. Warnings without a proposed correction were particularly weak: Table 3 implies that only 47.62% accurately identified an error. Ambiguity from the system should not be mistaken for authority.
Pedagogically, Grammarly is best positioned as an auxiliary second reader rather than a final judge. It can direct attention to a passage, but the learner must decide what is erroneous and which revision is better using grammar, context, and communicative purpose. Style feedback becomes more useful when transformed from an accept button into the question, “Does this expression still say what I intend?”
The study also teaches a broader lesson about AWE metrics. High precision among displayed flags can coexist with many undetected errors, and a strong overall average can conceal severe weakness in a specific category. Evaluation should therefore separate precision, recall, error-type performance, contextual range, and preservation of meaning after correction. The paper’s practical message is straightforward: automated feedback gains its value from immediacy, but turning that feedback into learning remains the responsibility of the writer and teacher.
댓글남기기