The hard question in AI before-after analysis for cosmetic surgery is not whether a model can make a face look more rested, narrower, younger, or smoother. Many can. The harder question is what that image is allowed to mean when it appears in a consultation: a visual thought experiment, a communication aid, a planning document, or an implied forecast of healing.
That distinction matters because cosmetic surgery outcomes are not only changes in contour. They include bruising, swelling, scar placement, skin texture, asymmetry, discoloration, delayed settling, and patient-specific healing. A generated “after” image can be attractive while removing exactly the features a surgeon needs to discuss before consent.

The first useful evidence asks surgeons, not just software
The most relevant evidence so far is Yassa et al.’s 2025 ePlasty study, described as the first peer-reviewed investigation comparing generative AI before-and-after facial aesthetic images with evaluation by board-certified surgeons. The study tested three platforms—Midjourney, Leonardo, and Stable Diffusion—across four procedure types: blepharoplasty, facelift, rhinoplasty, and brow lift. Six board-certified plastic surgeons rated the generated results across 11 dimensions, with prompts engineered through ChatGPT-4.[1]
That design is not large enough to settle the field. Six evaluators, three platforms, four facial procedures, and one prompt-engineering method leave plenty of room for model updates, prompt differences, specialty variation, and procedure-specific behavior. But it is the right kind of early question. The images were not judged only by whether they looked impressive to a general viewer; they were judged against dimensions that matter to clinical use.
| Study element | What was evaluated | Why it matters clinically |
|---|---|---|
| Procedures | Blepharoplasty, facelift, rhinoplasty, brow lift | These procedures involve different anatomy, scar patterns, swelling profiles, and aesthetic endpoints |
| Platforms | Midjourney, Leonardo, Stable Diffusion | The comparison tested commonly discussed generative image systems rather than a single vendor claim |
| Evaluators | Six board-certified plastic surgeons | Clinical realism was assessed by people accustomed to postoperative anatomy and patient counseling |
| Rating structure | Eleven evaluated dimensions | The study separated visual plausibility from features needed for clinical usefulness |
The separation between realism and usefulness is the hinge. A patient may reasonably react to a smooth after-image as “that looks like me, only better.” A surgeon has to ask different questions: Did the model alter only the requested anatomy? Did it respect likely incision behavior? Did it show plausible edema or bruising? Did it preserve skin texture? Would this image make counseling easier, or would it make a realistic explanation sound like backtracking?
Moderate realism is not the same as clinical value
The headline numbers are modest. Surgeon-rated realism scores for the AI-generated images ranged from 2.90 to 3.57 on a 5-point scale. Midjourney received a mean realism score of 3.57/5, compared with 2.90/5 for Stable Diffusion, a statistically significant difference in perceived realism (p < 0.01).[1]
If the question were only “which model makes the more convincing picture,” that result would matter more. In a clinical setting, it matters less than it first appears. The same study found no significant difference in clinical value between the models (p = 0.38).[1] In other words, the model that looked more realistic did not become clearly more useful for preoperative planning or counseling.
That is a familiar trap in aesthetic imaging. A more polished image can feel more authoritative, but polish is not anatomy. It may even be more dangerous when the image is almost plausible: good enough to be persuasive, not good enough to carry the clinical meaning attached to it.

The weakest point is the one patients most need to understand
The lowest-rated dimension across all 11 metrics was healing and scarring prediction, with scores significantly lower than most other evaluated metrics (p < 0.01).[1] This is not a peripheral flaw. Healing and scarring are central to the difference between a cosmetic rendering and a surgical result.
A generated rhinoplasty image may narrow a bridge or refine a tip. A generated blepharoplasty image may remove hooding and brighten the periorbital area. But postoperative reality also includes swelling that may obscure the final contour, bruising that changes the early appearance, scar maturation that takes time, and small asymmetries that may matter more to the patient than to the algorithm. If an after-image systematically underpredicts those features, it is not merely optimistic. It changes the preoperative conversation.
This is where a visually impressive AI image can make the clinician’s job harder. The surgeon who declines to promise a frictionless result may appear conservative or evasive when the screen is showing a clean, settled, scarless face. The image has already made a promise, even if no one says the word promise.
Failure modes that are small on-screen but large in clinic
Yassa et al. also reported evaluator concerns about the uncanny valley effect, unrealistic skin smoothness and texture, and unintended anatomical alterations beyond the described surgical procedure.[1] These errors are easy to underestimate if the image is treated as a consumer visual. In surgery, each one matters.
- Unrealistic skin texture can make a postoperative face look healed, poreless, and evenly colored before biology would permit it.
- The uncanny valley effect can create a face that is almost natural but subtly wrong, making it difficult for patients to identify what feels off.
- Unintended anatomical changes can imply benefits outside the planned procedure, such as a brow, nose, eyelid, or facial contour changing when it was not part of the surgical request.
- A flattering global edit can blur the boundary between surgical simulation, beautification filter, and identity alteration.
For a software team, those may look like artifacts to be smoothed out in the next model. For a surgeon, they are documentation problems, consent problems, and expectation-setting problems.
Why the consultation pressure is already real
The Guardian reported in May 2026 that plastic surgeons are increasingly being asked to create an “AI face,” with patients bringing AI-generated ideal images into consultations.[2] That report is journalism, not clinical evidence, and it should be read accordingly. Still, it identifies the practical pressure point: these images are not waiting for validation before entering the room.
Once a patient arrives with an AI-generated ideal, the surgeon must translate a synthetic face back into anatomy, tissue behavior, risk, and time. Some of that translation may be productive. An image can reveal priorities a patient struggles to verbalize: less upper-lid heaviness, a softer nasal tip, a more open brow. Used carefully, it can start a conversation.
The problem begins when the image is treated as a destination rather than a preference signal. A model that underrepresents bruising, swelling, scars, and ordinary imperfection does not just show an ideal. It can train the patient to expect the wrong category of result.
Broader AI accuracy claims do not solve this problem
Plastic surgery AI is an active research area, and some broader figures sound much more encouraging than the before-after image data. Arkoubi’s 2025 review and meta-analysis reported 25 included studies and an 88% pooled accuracy figure across AI applications in plastic and reconstructive surgery.[3] That number should not be imported into cosmetic before-after counseling as if it answers the same question.
The limitation is not that laboratory accuracy is irrelevant. It is that accuracy depends on the task. Classifying an image, predicting a measurement, supporting reconstruction workflow, or generating a plausible postoperative face are different problems. Arkoubi also noted that none of the 25 studies included prospective clinical trials or real-world deployment data.[3] Without that, pooled performance is context, not clearance for use in patient-facing surgical forecasting.
A separate comprehensive review of AI’s role in plastic surgery describes a wide range of possible applications across the specialty, from analysis and planning to workflow support.[4] That breadth is useful, but it also reinforces the need to avoid collapsing all AI tools into one category. A model may help with one bounded task and still be unsafe or misleading when asked to simulate an individual patient’s visible postoperative course.
Automated aesthetic assessment is adjacent, not equivalent
Varghaei et al.’s 2025 work on automated assessment of facial plastic surgery outcomes and the SurFace1259 dataset points to another adjacent direction: using public facial images to evaluate aesthetic outcomes computationally.[5] That line of work may become useful for measurement, benchmarking, or research support, but it carries its own constraints. The dataset was sourced from public Instagram posts, may include digitally retouched images, and lacks demographic metadata such as age, sex, and ethnicity, limiting generalizability.[5]
Those limitations matter because cosmetic outcome assessment is not demographically neutral and not immune to image culture. If the training or evaluation material is filtered, retouched, selectively posted, or demographically opaque, the resulting system may learn a public-facing aesthetic record rather than a clinically representative one.
What can be trusted today
The fairest answer is not that AI before-after images are useless. It is that they are easy to overuse. Based on the current evidence, they may be reasonable as exploratory images, preference elicitation tools, or conversation starters when the clinician explicitly frames them as non-predictive. They are not reliable enough to stand as surgical forecasts.
| Use case | Current judgment | Reason |
|---|---|---|
| Exploring aesthetic preferences | Potentially useful with careful framing | The image may help reveal what the patient likes or dislikes |
| Preoperative planning | Not suitable as a direct planning tool | The evidence does not show dependable procedure-specific clinical value |
| Patient counseling | Unsafe if presented as an expected result | Healing, scarring, bruising, swelling, and texture are underrepresented |
| Outcome documentation | Not appropriate | Generated images can introduce anatomy and appearance changes that did not occur |
| Marketing | High risk | A polished synthetic result can imply predictability that the evidence does not support |
No FDA-cleared or CE-marked AI tools specifically for cosmetic surgery before-after analysis were identified in the available search materials. That does not mean every use is prohibited, and it does not evaluate general imaging software or non-cosmetic AI tools. It does mean clinicians should not treat this as an adoption-ready device category with established regulatory and clinical validation behind it.
For now, the clinical standard should be plain. If an image cannot reliably represent healing, scarring, procedure-specific anatomy, and ordinary postoperative imperfection without adding impossible polish or unintended changes, it should not be treated as a forecast. It can sit in the consultation as a prompt. It should not sit there as evidence.
References
- Facial Aesthetics in Artificial Intelligence: First Investigation Comparing Results in a Generative AI Study — ePlasty, 2025.
- 'You can't control everything': the rise in plastic surgeons asked to create 'AI face' — The Guardian, May 2026.
- The Transformative Role of Artificial Intelligence in Plastic and Reconstructive Surgery: Challenges and Opportunities — 2025.
- The intelligent lift: AI's growing role in plastic surgery — a comprehensive review.
- Automated Assessment of Aesthetic Outcomes in Facial Plastic Surgery — Varghaei et al., 2025.
Comments
Join the discussion with an anonymous comment.