I Have Never Written a Blog Post Without AI. Including This One.

Table of Contents

That sentence makes some of my colleagues uncomfortable.

It should not. By the end of this post, I hope you will agree on one thing: if we do not decide, on purpose, what to hand to AI and what to keep, the decision gets made for us. Randomly. It already is.

Let me start somewhere far from academic writing. A hospital in Nepal.

👋🏽 Quick announcement: Learn how to use AI in research the right way in my 5-Week LIVE AI Academic Writing Accelerator (included with Research Boost subscription) • Starts September • Limited to 100 Seats → Reserve Your Seat HERE

1. The internship that ran on scut work

In medical school and internship in Nepal, a large part of my day was not medicine. It was logistics.

I drew the blood. I carried the sample to the lab and walked back for the report. When a patient needed a higher center, I rode in the ambulance, sometimes bagging and masking the whole way.

For a long time I believed that grind made me a doctor. I was wrong.

Something did build judgment in those years, but it was not the blood draws. It was the moments in between. Presenting a patient to a senior and being asked why. Guessing sick or not sick before the labs came back. Thousands of small calls, each corrected by someone better than me.

The scut work was the container. The loop inside it was the education. I got the loop by accident.

2. The loop without the scut

My residency in the United States was a different world. I never drew blood. I never pushed a stretcher.

Almost all of my training was judgment, taught on purpose:

  • Sick versus not sick.
  • Emergency versus can wait until morning.
  • The emotion the patient is not saying out loud.
  • How to speak so the patient actually hears you.

The method was explicit. Think out loud. Present your reasoning, not just your answer. Let the attending pull it apart. Present again tomorrow.

Same loop as Nepal. Zero scut. My judgment grew faster because the loop was the whole day, not the 10 minutes between errands.

Nobody should mourn the scut work. But residency here did something different: it decided what the intern’s time was for. That is the decision we are failing to make with AI.

Take the EHR. An AI summary of every rheumatology note in a chart is a real gift. But deciding what belongs in the history, what matters on exam, and how to reason through the assessment and plan is the skill the intern is there to build. Hand that to AI and the trainee gets faster today and worse for a decade.

That is not a hypothetical. In a 2025 study of experienced endoscopists, a few months of working with AI-assisted colonoscopy was followed by a drop in their detection rate when the AI was switched off, from about 28 percent to 22 percent. Skilled clinicians, measurably worse at the core task after routine AI exposure.

3. The rest of the knowledge economy just caught up

A recent newsletter piece describes what its author calls the “engineered apprenticeship.” I will call it the engineered internship, our word in medicine and research. Law, tax, audit, and consulting have all arrived, in the past year, at the lesson medicine learned a generation ago.

The argument is sharp. AI did not break training by making courses worse. It switched off the engine doing the real development: years of routine tasks, done near senior people whose reasoning rubbed off. AI took away the scut work. And with it, by accident, the loop hiding inside.

The response has been the same everywhere: Big Law deposition simulators, a KPMG tax simulator, Deloitte’s rebuilt audit program, PwC training first-years as reviewers of AI output from week one. Every build has three parts. A loop: make the call before you see the expert’s answer, compare, go again on something harder. A captured expert: someone senior narrating why they made the call, so reasoning gets taught instead of absorbed by seating chart. A target: judgment that used to arrive in year five, taught from day one.

4. My own test case: writing

I learned medical writing the traditional way (as most of you probably did). Case reports, then systematic reviews, then cohort studies. No course. I copied papers I admired and worked with whoever would have me. Years of reps, and almost nobody asking why I made the choices I made. It worked. It was slow.

Then AI arrived, and I started doing something new: writing for a general audience. Blog posts. LinkedIn. Newsletters.

Most of what I publish is AI-assisted in a big way. The research, the brainstorming, the structuring. If you took AI away tomorrow, I would struggle to write this post, because I never learned to do it the old way.

And yet I am not worried.

5. What I actually trained: judgment and taste

I cannot write 10 great hooks from a blank page. But when AI gives me 10, I know which one is great. The one that stops someone scrolling at 11 pm.

When a draft comes back, I can tell which paragraph is hollow, which claim I cannot stand behind, and which sentence does not sound like me.

That skill came from studying what great looks like: which posts do well on LinkedIn, what a real hook does. AI produces. I judge. I am accountable for what goes out under my name.

Judgment is knowing what good looks like before it exists. Taste is recognizing it the moment it shows up.

Production is cheap now. Selection is not.

6. The gap AI widens

On routine tasks, AI narrows the gap between novices and experts. In the best-known study, customer support agents given an AI assistant got 14 percent more productive on average, but the least experienced gained 34 percent while the most experienced gained almost nothing. On complex tasks that need domain judgment, it widens it. Experts use AI to go further. Novices accept output that is plausible but generic. When consultants at a top firm used GPT-4 on a task just outside what the model does well, they were 19 percentage points less likely to get the right answer than colleagues with no AI at all. The output read well. It was wrong.

Hand a fresh trainee an AI writing tool without teaching them to check it, and you do not close the gap between them and you. You compound it. They will produce a Discussion that reads smoothly and argues nothing, and they will not know the difference. A survey of knowledge workers found the pattern in miniature: the more people trusted the AI’s answer, the less they checked it.

That is the trainee who lost the loop along with the scut.

7. The real problem: nobody is deciding

In most labs right now, nobody has decided what trainees should outsource. So it happens by default. The research trainee uses AI for the literature search, which is fine. Then the first draft of the Introduction. Then the Discussion. Then interpreting the results. Then deciding what the paper is even claiming. Nobody chose that. It drifted, one convenient step at a time.

Random outsourcing does not stop at the tasks that are safe to give away. It keeps going until it reaches the ones that should never be given away, and it cannot tell the difference.

Residency solved this by being explicit. Scut work: outsourced. Presenting the assessment: never outsourced, practiced daily, corrected in public. Research training needs the same split, written down:

Outsource freely. Literature retrieval. Formatting. Reference management. First-pass summaries of papers you will read anyway. Grammar and flow on a draft whose argument you already own.

Keep for yourself, always. Interpreting your own data. The argument of the Discussion. Any number. Any citation you have not opened. The final read for voice. And the accountability for all of it.

Make trainees practice on purpose. What the paper is really claiming. Which result matters most. What a reviewer will attack first. When a plausible AI sentence is quietly wrong. These are the “sick versus not sick” of manuscript writing, and they cannot be learned by watching AI do them.

Your lists might differ from mine. Having them is what matters.

8. Running the loop

Once the lists exist, the loop is simple. Here is how I run it and how I am starting to run it with the people I mentor:

  1. Commit before you ask. Before AI drafts your Discussion, write three sentences on the main argument. Then compare. If AI’s version is better, ask why. That gap is the lesson. This is the testing effect in disguise: attempting an answer before seeing one makes the answer stick in a way that reading never does.
  2. Ask for 10, choose 1, explain the choice. 10 titles, 10 hooks, 10 framings of the limitation. Rank them and say why number 4 beats number 7. Ranking with reasons is judgment practice. Accepting the first output is not.
  3. Play the reviewer, not the author. Read an AI-drafted paragraph the way Reviewer 2 would.
  4. Make your mentor think out loud. When a senior author edits your draft, ask why they cut that sentence. Save the answers. Over a year they become your feedback rubric.
  5. Run AI rounds. Once a month, dissect a real AI-assisted draft as a group. Where was the AI confidently wrong? Where did the human catch it? This is M&M rounds for writing.
  6. Revisit the lists. Every few months, ask what has drifted. The task you outsourced “just this once” is now the default. Pull it back or make it official.

Every training program is making the outsourcing decision right now. And most are making it by not making it.

People can tell AI slop from expert-led work. What separates them is whether a trained judgment made the calls at the steps that matter. That judgment no longer shows up on its own. It has to be built, and building it starts with a list.

Top Papers on AI in research this week:

  1. Google’s AI Co-Scientist Goes Live – Google’s new multi-agent system, built on Gemini, ran real scientific experiments this month, not just literature summaries. It designed new material synthesis routes and helped uncover medical AI architectures in early trials. Expert reviewers say built-in safety checks cut down on hallucination and plagiarism along the way.
  2. Making Clinical AI Show Its Work – Clinical language models often ace hospital benchmarks, then quietly fail once deployed elsewhere. A new framework called CAST traces that failure to shortcut patterns the model leans on instead of real patient data. Suppressing those shortcuts during training kept accuracy high on mortality prediction while making the model’s reasoning something clinicians can actually check.
  3. AI Grades Research Designs, Imperfectly – A team of researchers built ARDTrA, a system that scans social science papers and scores how sound their research designs are. Passage length, not topic difficulty, explained most of the variation in how well it performed. And the papers that tripped up the AI were not the same ones that split human experts, hinting machines and people struggle for different reasons.
  4. Can AI Read Minds? – Researchers ran 2,099 AI instances across four model families through economic games designed to test whether they can read others’ intentions, the skill known as mentalization. GPT-5 matched or beat 251 human participants once given the right strategic prompts. The catch is that better prompting, not better reasoning on its own, drove most of the gain.

Top Papers on AI in education this week:

  1. The Campus AI Trust Gap – A survey of 2,121 people at one university found students using AI far more often than faculty or staff, and trusting it more too. Faculty and staff worried more about academic integrity instead. Clear institutional policy built more trust than any technical fix, the researchers found.
  2. A Tutor That Grades Its Own Explanations – A new intro-programming tutor called ESSE asks students to explain code in their own words, then grades those explanations the way a human expert would. Explanations grew more complete with each attempt, and students showed measurable learning gains. The system’s judgments held up against both expert and crowdsourced human graders.
  3. Comfort With AI Isn’t Understanding It – Estonian researchers built an 18-item test measuring what high schoolers actually understand about how generative AI works, then ran it on 7,432 students. Frequent use of AI tools and liking them told researchers almost nothing about whether students understood how the technology actually functions. Comfort with a tool, in other words, is not the same as understanding it.
  4. Teaching Engineers to Build With AI – The University of Arkansas built a three-level AI curriculum for mechanical engineering students, aimed squarely at gaps in thermal system modeling and cross-discipline collaboration. Students move through introductory, applied, and advanced projects rather than a single AI module bolted onto an existing course. The team published its syllabi and code so other engineering programs can copy the approach.
  5. Half of Students Already Use AI for Coursework – A survey of scholarship applicants found 54% openly used AI while writing their essays this year. Most used it for brainstorming, outlining, and editing, not for writing whole essays outright. Students still want a human making the final call in fields like healthcare, education, and law.

Leave a Comment

Your email address will not be published. Required fields are marked *

Related Posts

Join the ONLY NEWSLETTER You Need to Publish High-Impact Clinical Research Papers & Elevate Your Academic Career

I share proven systems for publishing high-impact clinical research using AI and open-access tools every Friday.