Sunday, 26 July 2026

VibeRounds: Early User Feedback on AI-Augmented Clinical Reasoning

 

VibeRounds: Early User Feedback on AI-Augmented Clinical Reasoning

A field report from the first weeks of testing

What is VibeRounds?

VibeRounds is a Socratic AI framework for clinical reasoning education, coined by Dr. Avinash Kumar Gupta in June 2026. The core idea is simple but deliberately unusual: instead of asking an AI to hand over a diagnosis, VibeRounds turns the AI into a Socratic attending — questioning a learner's reasoning, flagging cognitive biases, and withholding the full answer until the learner has committed to their own thinking first.

Built as a library of persona prompts and a larger "Clinical Cognition Operating System" (CCOS) — 57 reasoning modules, four pedagogical frameworks, and 200+ chainable pipelines — VibeRounds is designed to be pasted into any LLM (Claude, Gemini, ChatGPT) alongside a real or invented clinical case. The system enforces a few core constraints: learners must commit to an initial answer before receiving any hint, hints are tiered (framework → direction → partial answer), and low-effort replies get redirected rather than rewarded.

The philosophy is captured in the project's own tagline: "Not AI that answers — AI that questions."

Putting it to the test: a running feedback log

Over the course of about a week (20–27 June), a series of real and case-based clinical scenarios were run through three different LLMs — Claude, Gemini, and ChatGPT — using VibeRounds prompts. Below is a synthesis of what came back.

The good: reasoning over recall

The most consistent theme across entries is that VibeRounds succeeded at its central goal — shifting the interaction from "give me the answer" to "make me think." One early real patient-inspired case (20 Jun) was described as helping to structure clinical thinking, surface overlooked possibilities, and use gentle Socratic nudges instead of direct answers. A stroke case the next day was said to feel "like discussing with a professor," with questions that promoted critical thinking rather than recall and pushed the user to apply previously learned concepts to a real scenario.

This pattern repeated across multiple cases and models:

  • A ward-round style Gemini session (21 Jun) was praised for its stepwise reasoning, investigation selection, and management-focused discussion, closely resembling bedside teaching.
  • A Claude-run surgical case (23 Jun) made the user feel "like being a surgery resident" — immersive and enjoyable.
  • A diabetes-focused Gemini case (24 Jun) went beyond diagnosis into perioperative management, euglycemic DKA, and medication interactions, with the user noting that case-based learning this way "bridges textbook knowledge and patient care" and improves confidence.
  • A pulmonary embolism–anticoagulation case (27 Jun) was described as feeling "like solving a clinical puzzle," integrating physiology, pathology, and management across DKA, sepsis, ARDS, and PE — with Socratic questioning said to strengthen clinical reasoning "far better than memorizing isolated facts."

Several sessions also produced concrete, specific learning points users hadn't previously connected — sideroblastic anemia as an ATT complication, enoxaparin dosing and the Cockcroft–Gault formula with female correction factors, perioperative hypoglycemia and delirium management, and revisiting first-year anatomy in a clinical context. One user explicitly called a case their "favorite," citing improved patient-counselling skills as an unexpected bonus.

The mixed and critical: model and prompt sensitivity

Not every run landed. Feedback surfaced two consistent friction points:

1. Model matters as much as the prompt. A Gemini session (21 Jun) stopped after a single answer, possibly because a lighter model variant (Gemini Flash) was used — the user wanted the discussion to go "much deeper and longer." A separate Claude case the same day felt "somewhat odd," with the user wanting more cross-questioning and brainstorming than they received. Later, when a case-3 differential diagnosis exercise improved noticeably after switching from Claude to ChatGPT (23 Jun), the user themselves flagged the ambiguity: was the improvement due to the refined prompt, or simply the different LLM?

2. Prompt iteration made a visible difference. On 22 Jun, a user directly compared an older prompt version to a newer one on the same case type and preferred the newer one — citing a more complete walkthrough, histological differentiation, postoperative therapy discussion, and physiology review as more informative than the earlier version. This suggests the prompt engineering behind VibeRounds' personas is not incidental — refinements measurably changed the depth of the teaching output.

Some entries also came back with no substantive feedback at all (one Gemini case on 21 Jun), a reminder that engagement with this kind of open-ended Socratic tool varies session to session.

What this early log suggests

Across nine days and roughly fifteen logged sessions, three patterns stand out:

  1. The Socratic mechanism works when it's given room to work. Users repeatedly described the experience in terms of reasoning — differential-building, management logic, guideline lookup (e.g., needing to check Wells criteria for a PE case) — rather than passive answer retrieval.
  2. Output quality is not uniform across LLMs. The same persona prompt produced markedly different depth depending on which model — and which tier of that model — was used, with lighter/faster models sometimes truncating the Socratic exchange prematurely.
  3. Prompt versioning is a live variable. The 22 June comparison is early evidence that VibeRounds' prompt design is iterating in a direction users notice and prefer, though isolating "better prompt" from "better model" remains an open question the project will need cleaner A/B testing to resolve going forward.
More Details (Links to sessions by 4 medical students) - https://github.com/avi33tbtt/avi33tbtt.github.io/blob/master/demo/trials/trial3/week1.md

Where to learn more

VibeRounds is openly documented, CC BY 4.0 licensed, and built around a philosophy of "a paradigm, not a prescription" — meaning every persona and module is meant to be copied, adapted, and extended rather than used as-is. The full framework, including the CCOS module library, prompt builder, and courseware, is available at avi33tbtt.github.io.


This article is based on an internal user feedback log spanning 20–27 June, covering real and simulated clinical cases run across Claude, Gemini, and ChatGPT.

No comments:

Post a Comment