The novelist and columnist John Paul Brammer was asked recently how he can tell when prose was written by AI. His answer:
A certain appetite for a certain style. “There’s this corny animism everywhere. The bench that watches. The rock that remembers.”
Anyone who reads seriously knows exactly what he means. He said it after a debut crime novel that reportedly sold for $2 million was withdrawn by the author’s own agents, who said they could no longer authenticate how the manuscript evolved from origin to completion.
So the industry is reading for tells, and in fiction that is completely legitimate. A novel is its sentences. A reader who closes the book because the prose rings hollow has judged the whole deliverable.
Cross into science, and the same reflex arrives intact, aimed at something it was never built to measure. A paper is a claim that either holds up under its design and data or does not. The writing decides whether that claim gets understood. It does not decide whether it is true.
You know how this goes. Second paragraph of your introduction, the reviewer hits “delves into.” Then “underscores the crucial importance.” Then a tidy three-part list where two items would have done. Something shifts. They have not decided your randomization was flawed. They have stopped reading as a colleague and started reading as an inspector. By the time they reach your limitations, they have an opinion about your study.
👋🏽 Quick announcement: I would love to have you in my FREE 5-Week LIVE AI Academic Writing Accelerator • Starts September • Only 100 Seats → Apply for Your FREE Seat
Writing has always counted. This is not that.
There is an obvious objection here, and it is half right.
Science has never judged the work alone. Manuscripts and grants have always been evaluated partly on how well the science is communicated:
Entire courses on grantsmanship. Workshops on manuscript structure.
I have spent years teaching early-career researchers that story, structure, and precision decide whether their work gets read and cited.
So I am the last person who will tell you the words do not matter.
But look at what that tradition evaluates. Clarity. Precision. A line of logic you can follow from aim to conclusion. When a reviewer downgrades a grant because the specific aims are muddled, the writing is telling them something real about the science. An applicant who cannot state the hypothesis cleanly may not have a clean hypothesis.
The AI-ism test judges writing for something else entirely. Provenance. What tool may have touched it.

“Delves” is not unclear. An em dash is not imprecise. A sentence can contain both and communicate its science exactly. When a reviewer recoils at those markers, they have not caught a communication failure. What they believe they caught is a fingerprint, and that fingerprint is made of style almost by definition. Across 15 million PubMed abstracts, Kobak and colleagues found 379 excess words in 2024, 66% of them verbs and 14% adjectives. During COVID the surge words were respiratory and remdesivir. After ChatGPT, delves alone ran at 28 times its former rate.
So here is the question I want asked out loud: If the science is communicated clearly and precisely, should a single verb create backlash against the science itself?
Polished help was never disqualifying before. Departmental editors, writing centers, the mentor who rewrote your discussion line by line. Nobody ever argued that a manuscript improved by an editor invalidated the data. The provenance of the prose was not the test. The clarity was.
Now the two pull in opposite directions. Under the old test you revise toward clarity. Under the new one, researchers revise away from it: deleting em dashes they have used their whole careers, roughing up clean sentences, avoiding precise words that landed on a list. We spent decades teaching scientists to write better. We are now teaching them to write worse on purpose, so they can pass as human. (If you have ever used an AI humanizer, you know exactly what I mean.)
What I keep hearing in the hallway
I was on an NIH Grant Awardees Forum recently, and the officials were relaxed: use the tools, as long as the ideas are yours and the science belongs to you. NIH’s own notice, though vague, agrees. A rule against outsourcing the science, not against using the tool.
Then I talked to the researchers, and I felt like I was in a different room.
- One had been stuck on a project, and a family member suggested they ask an AI for research ideas. They felt deeply insulted. Something they had spent years training to do, and a chatbot was being floated as a substitute. How dare anyone even suggest that?
- Another was uncomfortable that their mentor had put their research into an AI to get feedback. It was the institution’s own secure AI platform. It still felt wrong to them. Something they had built with their own hands, handed to a machine without anyone asking.
- A senior mentor was angry and disappointed because their mentee had looked up a method in ChatGPT, and ChatGPT invented things. They had expected more of the mentee.
That last one is worth sitting with. The mentee’s failure was not opening ChatGPT. It was not verifying the output, and nobody had taught them to verify, because in a group where the tool is shameful you learn someone is using it at the moment it breaks. Stigma feels like a safety control. It guarantees you find out last.
None of this is what researchers say when nobody is watching. In Nature’s poll of 5,000 researchers, more than 90% were comfortable using AI to edit or translate their writing. Privately the field has made up its mind. Publicly we still perform suspicion for each other.
Some version of that room happens at nearly every talk I give. The last one, at the Rheumatology Research Foundation, came back rated the best talk of the session and the worst talk of the session. Two audiences, the same 40 minutes, opposite verdicts on whether I should have been allowed to give it.
Somewhere in the last 2 years, AI in research unfortunately became like religion or politics at a dinner table.
Those who get hit hardest are the ones who need the help
I came into research as an international medical graduate with no research background and zero publications. English is not my first language.
My early manuscripts were very careful. Slow, plain, controlled. I would rewrite a sentence 6 times until I was certain it could not be misread, because I did not trust my instincts in the language.
That is precisely the kind of writing AI detectors flag.
Liang and colleagues at Stanford ran seven detectors across 91 TOEFL essays by non-native English speakers and found an average false positive rate of 61%, against 5% for essays by American eighth graders.
The principal mechanism of AI detection is low lexical variability. Careful, controlled prose reads as synthetic.
And it lands on people already paying a toll. Non-native speakers are 2.5 times more likely to have a paper rejected on language grounds, and 12.5 times more likely to be asked to improve their English on revision. AI is the first tool that meaningfully narrows that gap. A detection culture punishes them for reaching for it. Every international graduate and ESL postdoc in your department is sitting in that position right now.
Humans do no better than the tools. Given a mix of human-written and AI-generated abstracts, 17 experienced reviewers in “The Great AI Witch Hunt” rated both as equally likely to be AI, and named opposite features as proof of the same conclusion. Convoluted sentences were an AI marker for 13% of them and a human marker for 3%. Write badly, suspicious. Write well, suspicious.
Where I stand
I disclose AI use in every manuscript and every grant. Every platform assures me that will not bias the review but I’m not sure how much I believe that. Across 13 experiments with more than 5,000 participants, professors and analysts among the evaluators, Schilke and Reimann found disclosure consistently lowered trust.
I disclose anyway, partly because it is our obligation and partly because of one finding in that same work. Trust fell furthest of all when someone did not disclose and a third party surfaced it later.
So I keep dated drafts and verify every citation against a source I opened myself. The question, the design, and the interpretation never come from a model. And I still cannot tell you whether disclosing protects me or marks me.
3 things we should strive for
- Judge the writing for clarity, not provenance. If the prose obscures the science, say so. That critique is fair and it gives the author something to fix. If it merely sounds like a machine touched it, you have a feeling about someone’s diction, not a finding.
- Never put an AI suspicion in a review without evidence that is not stylistic. A fabricated citation counts. A method that does not exist counts. A verb does not.
- Refuse AI detector scores as evidence, in your department, on your editorial board, in your promotion committee. The case that publishers should not use them predates this panic: the tools are unreliable, and the accusations they generate damage real careers.
In fiction, the prose is the work, and a critic who judges it is doing the job properly. In science, the prose is the lens the work is read through. Check that it is clean and in focus. That is what grantsmanship has always meant. What we are doing instead is dusting it for prints.
Have you ever been suspected, or come close? Have you ever flagged a manuscript because of how it sounded rather than what it showed?
Top Papers on AI in research this week
- Pathology Foundation Model That Talks Back – A 4.6-billion-parameter model trained on 2.3 million whole slide images and 700,000 pathology reports hit 0.967 AUC for pan-cancer detection. On rare cancers it barely slipped, to 0.957. It matched commercial clinical-grade products without any task-specific training.
- AI Titrates Oxygen Better Than Manual Care – A four-hospital randomized trial put 300 adults on either automated oxygen control or standard bedside titration. The AI group stayed in the target saturation range 85% of the time, versus 63% for usual care. Both hypoxemia and hyperoxemia fell, with no increase in serious adverse events.
- Explanations Help Doctors and Hurt Everyone Else – MIT researchers tested explainable AI on 153 primary care physicians and 623 lay participants making skin lesion calls. Physician accuracy jumped 21.5 points, and doctors pushed back when the model was wrong. Lay users did the opposite. When the AI erred, their accuracy dropped 21.1%.
- AI Agents Are Auditing the Literature – SAI Labs turned agents loose on 168 ICML 2026 papers to check whether the claims actually held. Of the 92 papers with testable claims, only 8 had more than 80% replicate. A separate Stanford analysis found errors per NeurIPS paper climbed 55% between 2021 and 2025.
- Sleep Studies Hold Signal We Have Been Discarding – Cleveland Clinic trained a model on raw polysomnography instead of summary scores. It sorted patients into five risk tiers. The top tier carried twice the five-year mortality risk of the bottom, a gradient the standard apnea-hypopnea index misses entirely.
- Who Actually Allows AI in Peer Review – Researchers catalogued reviewer AI policies at 111 conferences and journals. Only 21% of AI and NLP venues prohibit AI-assisted reviewing, against 58% of medical journals. They then scored machine reviews against human ones. GPT-5 tracked human scores at r=0.62, but AI reviews averaged 6.8 to 7.9 where humans gave 4.3.
- The Benchmark Was Broken, Not the Models – Domain experts audited all 65 problems in SciCode, a scientific coding benchmark, and found 263 defects. Most of them wrongly rejected correct answers. After repair, main-problem accuracy across twelve frontier models rose from 9-27% to 69-92%. The much-discussed 2026 plateau turns out to be an instrument artifact.
- Does AI Really Make Researchers More Productive – A critique had argued the observed productivity gain from LLM adoption was a mirage created by how adoption dates get assigned. The original authors re-ran the analysis five different ways. Output still rose roughly 17% among adopters, while pre-ChatGPT placebo tests came back flat at -0.9%.
- Agents That Run Meta-Analyses End to End – This framework converts structured evidence into runnable meta-analysis code, with deterministic validators checking the model at every step. Across 58 synthesis units, 98.2% of confidence intervals overlapped the published findings. Against direct LLM generation, it won on synthesis structure 57 times out of 58.
- More Papers, Worse Papers – A modelling study borrowed optimal foraging theory from ecology to ask what LLMs do to the research process. Even assuming the tools are cheap, fast, and accurate, it predicts more output at lower quality. The reason is incentives. Nobody spends the saved time on discretionary refinement when the reward is volume.
Top Papers on AI in education this week
- Can an AI Tutor Hold a Relationship for 30 Days – Most tutoring benchmarks test a single exchange. This one runs an agent through a month of sessions with a simulated learner whose knowledge state is grounded in real student data. Across 55 scenarios and ten agent configurations, almost nothing sustained good teaching over the full horizon.
- Teaching a Small Model to Tutor Like a Teacher – Researchers annotated 260 authentic teacher-student conversations with 32,379 labels covering thirteen tutoring strategies. That corpus was used to post-train a 4-billion-parameter model. It beat its own backbone by 20.3 points, outscored proprietary models, and took the top rating in a blinded study with 50 learners.
- Stopping Tutors From Giving Away the Answer – Pedagogical leakage gets a formal definition here: revealing the solution before the student has earned it. The author built a mediation layer with explicit disclosure contracts and an auditable release gate. Across 599 tutor responses, leakage flags fell from 181 to zero. Helpfulness ratings fell too, which is the honest tradeoff.
- LLM Guardrails Fail Hardest in Classrooms – A new K-12 safety framework tested ten models across 28 education-specific risk subcategories. Single-turn attacks succeeded 29.9% of the time. Sustained multi-turn attacks succeeded 53.6% of the time. Academic misconduct was the weakest category by a wide margin, at 80.3%.
- Universities Are Giving Up on AI Detectors – Yale, Vanderbilt, Johns Hopkins, and Indiana have banned or discouraged detection tools, and at least a dozen more institutions have switched off Turnitin’s. False positives and bias against non-native English writers drove the shift. Meanwhile 95% of faculty report fearing student overreliance, and 73% have handled an integrity case personally.
- Turnitin Wants Your Students’ Keystrokes – The company launched Clarity, which logs keystrokes, tracks revisions and pastes, and times writing sessions. Instructors get a replay of how a document came to exist. Its text detector was 70-80% accurate in a 2023 study. Retyping AI output defeats the entire mechanism.
- Some Faculty Are Simply Leaving – Job exits among educators over 55 rose from roughly 11% to 16% after ChatGPT launched. Professors profiled in the piece moved retirement forward by three to nine years. One instructor estimates that 150 of her 320 students use AI to cheat.
- Medical Students Use AI and Distrust It at Once – A survey of 358 medical students found 88.5% had already used AI tools. Yet 41.1% believed it could erode their clinical reasoning, and 59.5% flagged ethical or legal concerns. Digital literacy was the strongest predictor of positive attitudes. Prior use predicted nothing once adjusted.
- Treat Every Prompt Like an Experiment – A Vanderbilt scientist lays out a ten-point protocol for prompting with the rigor you would bring to a bench experiment. His own cohort data makes the case. Among incoming biomedical PhD students, 81% had used AI for science but only 5% could write a competent prompt. After instruction, 48% could.
- LLMs Pass the Bar and Fail the Notary Exam – Italian researchers had frontier models write full exam papers for the bar, judicial, and notary examinations. Submissions were anonymized among human papers and graded blind by real examiners. Gemini scored 79 out of 100 on the bar against 62 for its human comparator. Every model failed the notary exam, which demands goal-directed planning under strict formal constraints.
