Writing practice research · October 2026

Writing-practice research methodology: October 2026 dataset and reporting rules

This series analyzes 2,370 recent practice responses from 1,383 accounts. It publishes aggregate tables with explicit duplicate handling, contributor limits and score-interpretation boundaries.

Published · Research publisher: Lucas Weaver

Retained responses
2,370
Publication
5 October 2026
Evidence
AI practice estimates

Submission window: 22 August–5 October 2026. Scores are stored AI practice estimates. This self-selected sample does not establish official results, scoring accuracy or causal learning gains.

01

What does the October dataset contain?

The research window begins on 22 August 2026 UTC and ends at the frozen database snapshot on 5 October 2026. The source recorded 2,842 completed AI evaluation rows attached to submissions in that window. The final corpus contains 2,370 retained responses after the rules below.

The source is the shared database serving IELTS Writing Checker and Cambridge Writing Checker. ScoreQwik publishes the cross-app methodological reports but its imported copies are excluded. PTE has no qualifying responses in this recent window, so this release does not extend the historical PTE pilot.

The 22 August start is deliberately later than the date shown on the updated research privacy notices. Earlier submissions are outside this new analysis. The August benchmarks remain distinct publications with their own cohorts; this series is not a direct numerical update to their score findings.

Research corpus preparation: retained counts at each stage
StageRows or responses
Completed AI evaluation rows in the date window2,842
Latest completed evaluation per submission2,842
Supported site, exam and task labels2,842
Valid scores, complete criteria and 20–3,000 words2,842
After normalized exact-text deduplication2,672
After high-frequency-account exclusion2,370
Responses in cohorts qualifying for detailed reports2,087

02

How are evaluations and duplicate responses selected?

For each submission we select the latest completed AI evaluation created before the snapshot, breaking equal timestamps by evaluation ID. We then require the intended checker site, a supported task label, a numeric overall estimate and all four numeric criterion estimates within the native scale: 0–9 for IELTS and 0–5 for Cambridge.

Response text is lowercased, whitespace runs are collapsed and outer whitespace is removed inside the database. We hash the normalized text and retain the earliest instance within the same exam family, exam variant and task. Equal creation times are resolved by submission ID. Hashes and response text are not exported.

This removes exact normalized copies. It does not identify paraphrases, lightly edited templates, generated answers or the same response submitted under another task label. The term unique therefore means unique under this stated rule, rather than guaranteed independent or independently authored writing.

We exclude accounts contributing more than 50 retained distinct responses across the study window. This exploratory rule was chosen during the audit. It is a conservative eligibility decision, not proof that the excluded writing was invalid.

03

How are word counts, score profiles and correlations calculated?

Word counts are recomputed as whitespace-delimited words after whitespace normalization. In 27 retained responses (1.1%), this count differs from the stored product count. The research uses one consistent counting convention; it does not claim to reproduce official examiner word-count rules.

Responses need 20–3,000 whitespace-delimited words for this series. We do not apply official task minima as a research filter, because doing so would truncate length analyses. Nevertheless, product submission checks mean the recent evaluated data has little or no coverage below those minima.

Medians and quartiles use continuous percentiles. Means describe the stored estimate scale and should not be interpreted as precise amounts of underlying language ability. An overall estimate is the stored overall output, rather than a replacement calculated from its four criteria.

Length–score association uses Spearman correlation: Pearson correlation of average ranks, with ties assigned the midpoint of their rank positions. We provide descriptive coefficients, not population inference, p-values or evidence that changing length causes a score change.

A criterion shares the lowest estimate when it equals the minimum within the response. It is uniquely lowest only when no other criterion has that value. The mean gap to highest compares it with the maximum criterion estimate in the same response.

04

Why are some cohorts and table cells withheld?

Detailed primary cohort results require at least 100 responses and 30 contributing accounts. A first-per-account sensitivity sample can be smaller, but it belongs to an eligible primary cohort and must contain at least 30 accounts. Detailed length and score bins require 20 responses and 10 accounts.

Five cohorts qualify for detailed reports, representing 2,087 responses. The other 283 retained responses remain in the corpus inventory without detailed score or correlation tables. They should not be treated as missing evaluations or zero-score responses.

Suppressed score tails and length groups are omitted from downloadable detailed tables as well as from the page. Percentages use the full eligible cohort, so displayed bins may sum to less than 100%. Account counts across tasks or bins can overlap and should not be added to infer distinct people.

Primary cohorts eligible for detailed reporting
Task cohortResponsesContributing accounts
IELTS · Academic · Writing Task 1677372
IELTS · Academic · Writing Task 2923546
Cambridge · B2 First · Essay134129
Cambridge · C1 Advanced · Essay230212
IELTS · General · Writing Task 212351

05

What can the evidence establish, and where does it stop?

These observations describe people who chose to use an online writing checker. They are not a probability sample of examination candidates, and the site does not verify timed conditions, independent authorship or external assistance. Selected task level is not verified learner proficiency.

The evaluation rows do not record a complete model-and-prompt version history. Numerical comparisons therefore describe historical practice estimates inside a defined recent cohort. They do not validate official bands, measure exam pass rates, establish criterion difficulty or demonstrate learning gains.

No blinded examiner study, independent error coding or human pedagogical review was conducted for this release. Automated computation and drafting assisted the preparation of the reports. The publication identifies Lucas Weaver as the research publisher; it does not claim journal peer review.

Public downloads contain aggregates only. No essay text, task prompt, feedback quotation, account identifier, email or normalized-text hash is included. Readers receive the query and immutable published aggregate files, rather than access to the private database.

The query version is 2026-10-05.1; its SHA-256 is 6d1fac9b79df8d875a566c60cf1a9fe218593771cf652bb0eacb4ae70ca69980. Frozen ID ceilings and the snapshot timestamp bound the extraction. Because production rows can later be changed or deleted, rerunning the query is not a promise of an identical historical database state. The saved downloadable aggregates are the exact evidence for this publication.

Methods

How this report was prepared

The study uses the latest completed AI evaluation for each eligible submission in the 22 August–5 October 2026 window. Site, task, score and word-count checks precede normalized exact-text deduplication. Accounts contributing more than 50 distinct retained responses in the window are excluded by an exploratory audit rule.

Detailed primary cohorts require 100 responses and 30 accounts. Length and score bins require 20 responses and 10 accounts. The first-response-per-account-and-task analysis is a sensitivity check; it can contain fewer responses than the primary cohort. Account counts across tasks and bins can overlap.

The source does not identify every evaluator model and prompt version, verify examination conditions, or provide independent examiner scores. Lightly edited duplicates and external assistance may remain. Differences describe this practice sample and scoring system; they do not demonstrate official proficiency, causal improvement or universal task difficulty.

Computation and drafting were assisted by AI. This release has no independent human examiner validation, pedagogical review or journal peer review. Only aggregates are published; raw writing, prompts, feedback and learner identifiers are excluded.

Questions about the findings

Can these practice estimates be converted into official exam outcomes?

No. This series has no validated examiner-agreement or exam-outcome dataset. Each report keeps its native practice-estimate scale.

Does the October window demonstrate improvement since August?

No. The windows, contributor mix and cleaning decisions differ. The series makes no longitudinal improvement claim.

Evidence

Aggregate data, sources and citation

The downloadable files contain the report’s eligible cohort summaries, criterion profiles and unsuppressed length and score bins. CSV uses one row per measure; JSON preserves the table groupings. Neither file contains raw essays or learner identifiers.

Dataset version and integrity

Version 2026-10-05.1. SHA-256 of this report’s JSON file:

38a4b3fe80c4bcb7b95e0a514207df2c045386aa96318a1ed533c550112051d8

Official and contextual sources

Suggested citation

Weaver, Lucas. “Writing-practice research methodology: October 2026 dataset and reporting rules.” ScoreQwik, 5 October 2026. Version 2026-10-05.1. https://scoreqwik.com/research/writing-practice-research-methodology-october-2026

ScoreQwik is an independent practice service. These reports are not affiliated with or endorsed by the examination organizations.