| **Frontiers in Cell and Developmental Biology, February 13–16, 2024 | A peer-reviewed paper, disclosed AI-generated figures, a reviewer warning, and a publishing process that did not ensure the warning was resolved.** |
A note before reading: This case is sometimes told as “AI made up an image and no one noticed.” The public record shows something more instructive. The manuscript identified Midjourney in a figure caption, and Frontiers later reported that one reviewer raised concerns about the figures and requested revisions. The authors did not respond to those requests, yet the article was published. The case therefore concerns more than disclosure or individual attentiveness: it asks how a workflow records a warning, assigns responsibility for resolving it, and prevents publication while a required response remains outstanding.
On February 13, 2024, the journal Frontiers in Cell and Developmental Biology published a review article on the role of JAK/STAT signaling pathways in spermatogonial stem cells. The article listed three authors across two hospital affiliations in Xi’an, China, two reviewers, and a handling editor.
The paper contained an illustrative anatomical figure of a rat. Its caption stated that the images in the figure were generated using Midjourney, an AI image-generation tool.
The figure was not anatomically possible. The rat was depicted with grossly oversized genitalia rendered in a configuration no actual rat possesses. The labels surrounding the figure were nonsense words formed by AI hallucination: “dissilced,” “Testtomcels,” “Stemm cells,” “iollotte sserotgomar cell.” The labels did not correspond to any real anatomical structures or biological terms.
The article moved through four broad stages. The available sources establish some events but do not provide a complete record of every participant’s actions or reasoning.
1. Author preparation and submission. Three authors were listed, and the submitted manuscript contained the AI-generated figures and the Midjourney caption. The public record does not identify who prompted, selected, labelled, or inserted the images; how the authors divided figure-review responsibilities; or whether all three inspected the final figures before submission.
2. Peer review. Two reviewers were named with the article. Frontiers’ official post-publication statement says that one reviewer raised valid concerns about the figures and requested revisions. The statement does not specify every concern, reproduce the full exchange, or explain the other reviewer’s assessment.
3. Editorial decision and workflow control. Frontiers says that the authors failed to respond to the reviewer’s requests. Nevertheless, the article proceeded to publication. Frontiers subsequently investigated how its processes failed to act on the authors’ lack of compliance. The public statements do not establish the handling editor’s actions, the internal status of the requested revisions, or the precise control failure that allowed the article to advance.
4. Production and publication. The article, including the figures, was published online on February 13. The available record does not show which scientific-content checks, if any, were expected during production. Production is therefore a workflow stage to examine, not a group to blame without evidence.
After publication, images from the article circulated online and readers raised concerns about the figures. Frontiers credited community feedback with bringing the problem to its attention.
Frontiers retracted the paper on February 16, 2024. The formal notice said that concerns had been raised about the AI-generated figures and that the article did not meet the journal’s standards of editorial and scientific rigor. A separate Frontiers statement reported the unanswered reviewer request and said the publisher was investigating why its processes failed to act.
The public record does not provide a complete account from the authors, both reviewers, the handling editor, and production staff. Claims about what each person noticed, assumed, or expected should therefore be treated as hypotheses rather than facts.
The reason this case is teachable is that the AI-generated figure was visibly defective, the manuscript disclosed use of Midjourney, and at least one reviewer raised concerns—yet the workflow still did not stop publication. A more plausible-looking but inaccurate image could be harder to detect, which makes reliable follow-through even more important.
The teamwork question is therefore specific:
Disclosure alone did not prevent the failure. Nor was the problem simply a total absence of human scrutiny: a warning existed. What failed was accountable follow-through on that warning.
A functioning verification and escalation protocol would specify what must be checked, who must respond, who decides whether the response is adequate, how unresolved concerns are recorded, and what blocks the work from advancing. The case shows that a reviewer’s request has little protective value if the workflow does not ensure or enforce a response before publication.
Distributed responsibility may have contributed, but the sources do not reveal participants’ mental states or justify claiming that each person deferred to someone else. Students can use responsibility diffusion as one possible explanation, compare it with alternatives such as a status-tracking or decision-control failure, and identify what additional evidence would be needed to distinguish among them.
Disclosure versus verification. The caption identified Midjourney, but that disclosure did not ensure scientific validity. What is the difference between disclosing a tool and verifying its output? What must a team record about the use, review, and approval of AI-generated material?
Warning without closure. At least one reviewer raised concerns and requested revisions, but the authors did not respond and the article was still published. What workflow controls should make an unresolved reviewer request visible and publication-blocking? Who should have authority to close the concern?
Facts, hypotheses, and accountability. We know that a warning existed, but we do not know what every participant saw or assumed. Which explanations are supported by the record, which are hypotheses, and what additional evidence would you seek? How can a team assign accountability without inventing individual motives?
AI as a category requiring different protocols. What would justify treating AI-generated outputs as requiring different verification from other outputs? Should the protocol depend on the tool, the type of claim, the consequences of error, or the output’s verifiability?
Apply this to your team. Consider a moment when a teammate uses AI for part of a deliverable—a draft section, literature search, data analysis, or image. What is your team’s verification and escalation protocol? What gets checked, by whom, where are concerns recorded, and what prevents submission until they are resolved?