Connect with us

NEWS

AI Ghost Authors Carry Real DOIs Into Scholarly Databases

A June study found 1,655 Zenodo records listing AI-invented authors with real DataCite DOIs, including 991 registered in March 2026.

Published

on

A June study found 1,655 Zenodo records that list AI-invented authors and carry real DataCite DOIs. 991 of those records were registered in March 2026.

The names are a training fingerprint. The damage is that CERN’s open repository minted identifiers that Google Scholar and Semantic Scholar can ingest as if they named living scientists.

Real DataCite DOIs on Papers Nobody Wrote

Michał Brzozowski and Neo Christopher Chung, computer scientists at the Samsung AI Center in Warsaw, posted the preprint on 1 June 2026 and revised it on 29 July. They did not start with a hunt for fake journals. They started with a model test that kept coughing up the same people.

When they searched the open web and then scholarly repositories, those people showed up as volcano experts, astronauts, podcast hosts, and academic co-authors. On Zenodo, which CERN operates and which mints identifiers in the 10.5281/zenodo series, they found 1,655 ghost-authored records on Zenodo that claimed journals which do not exist and listed publication dates the server itself contradicts.

Zenodo registers those identifiers immediately with DataCite. DOI names provided by DataCite are built to persist, and any aggregator that harvests DOI metadata can pull the record. The authors write that Google Scholar and Semantic Scholar already index this material without checking that the named researchers exist in the fields they claim.

A DOI proves that metadata was deposited. It does not prove that anyone reviewed the file, that the journal exists, or that the author is a person.

Claude Kept Reaching for Elena Vasquez

The tell is not a single common name. Real people named Elena Vasquez and Marcus Chen exist. The authors’ claim is narrower: the expertise, the affiliation, and the name show up together only in AI-generated pages.

They probed nine Claude versions, ten GPT versions, and one Gemini model, all released between 2024 and 2026. Each checkpoint got 30 single-expert prompts and 30 pair prompts, temperature 1.0, in March 2026. Claude did not draw names independently. It produced a small ensemble.

MODEL NAME FINGERPRINTS

Model family Recurring names Test result
Claude (Sonnet 4, May 2025) Elena Vasquez, Marcus Chen, Amara Okafor Pair in 23% of pair prompts; Vasquez in 66% of single prompts
Gemini (2.5 Flash) Aris Thorne, Lena Petrova Pair in 37% of tests; Thorne in 93% of that checkpoint’s name draws
GPT (several versions) Elara Voss Strong solo name, no fixed partner

On Claude Sonnet 4, the Vasquez and Chen pairing appeared together in about 23% of pair prompts. By Claude Sonnet 4.6, that pairing had dropped out of the sample. An earlier Claude default, Elena Rodriguez, was already gone from every checkpoint they tested by October 2025. The names are version-specific, which makes the web a dated archive of which model wrote the page.

Gemini’s pair is Aris Thorne and Lena Petrova. GPT keeps returning Elara Voss and does not lock her to one co-author. Hybrid strings also appear, including Lena Voss, which mixes a Gemini first name with the GPT surname.

If you see Elena Vasquez online, she might be real. But if you see Elena Vasquez AND Marcus Chen together? Claude made them up.

Michał Brzozowski, Samsung AI Center Warsaw, on X, 15 June 2026

991 Records Landed in a Single Month

The Zenodo set is the part that leaves the open web and enters the scholarly record. The authors collected a corpus without querying “Elena Vasquez” by name, and she still ranked as the most frequent author in it. Server-side DataCite timestamps, which depositors do not control, show the files were not posted when the records claim.

Through 2025 the same collection shows a baseline of one to two records a month. Then the pace breaks.

THE ZENODO UPLOAD BURST

  1. Through 2025: One to two records a month in the collection.
  2. March 2026: 991 records registered in a single month.
  3. April 2026: 666 further records, about 25 a day across a 60-day stretch with March.
  4. 1 June 2026: The Ghost Couple preprint is posted on arXiv; a second version follows on 29 July 2026.

Those monthly counts (991 and 666) describe the burst of uploads. The 1,655 figure is the ghost-authored set the authors identified, a slightly tighter group than the raw two-month dump. No lab group posts at 25 records a day for two months. The timestamps also show the listed publication dates were pushed back by years, which is the difference between a new dump and a fake back catalogue.

Sidney Wong, a computational linguist at the University of Otago in Dunedin, New Zealand, called the scale of the operation “scary”. The names make the fakes searchable. The DOIs make them durable.

Mei-Lin Zhang’s 35 Papers Span Five Fields

ResearchGate is the other shelf. Ghost names there form groups whose co-authors come from more than one model family, which is not how a real lab cites its own people. Marcus Chen appears on two ResearchGate papers whose co-author lists lean on other AI-typical names, including Anya Sharma, Elena Rodriguez, and Anika Sharma.

The most structured case is a profile for Mei-Lin Zhang, presented as a researcher at Covenant University in Ota, Nigeria, with an institutional email check on the page. As of 15 May 2026 the profile listed 35 publications, 0 citations, and 420 reads. Zhang sits in last-author position on the set. The papers were uploaded from August 2023 through April 2026, and they do not form a research programme.

FIELDS ON ONE PROFILE

  • Kubernetes security: Cluster defence papers sit beside work that has no computing overlap.
  • Oil refinery decarbonization: Energy-process claims share an author list with farm-extension studies.
  • Agricultural extension: One October 2024 paper, Linking Farmers to Markets, names Aris Thorne and Elena Vasquez as co-authors with Samuel P Okonkwo and Mei-Lin Zhang.
  • IoT cryptography: Device-security titles continue the same last-author pattern.
  • Radiographic imaging: Clinical imaging sits in the same 35-paper stack, still with 0 citations.

Elena Vasquez, a Claude default, appears on at least three of those papers. Aris Thorne is a Gemini default. Putting both on Linking Farmers to Markets is cross-model ghost co-authorship on a single fabricated paper. Covenant University is named by the profile; the preprint does not show the university standing behind the person.

Publication dates on these ResearchGate records also track model release windows, which gives the authors a clock: when a ghost pair enters the site, a given checkpoint was in the wild.

OpenAI’s Goblin Problem Is the Same Kind of Quirk

Brzozowski does not claim to know why Claude locked onto Vasquez and Chen. He points to a documented training accident at OpenAI. Some ChatGPT versions started calling software bugs gremlins and dragging goblins into answers that had nothing to do with folklore.

OpenAI tied the habit to a “nerdy” and “playful” personality setting. During training, the model was rewarded for creature metaphors, and the tic spread into general models. After GPT-5.1 launched in November 2025, OpenAI recorded a 175% rise in the word goblin and a 52% rise in gremlin. The company later removed the nerdy personality and added system instructions. The wording still appeared at times in GPT-5.5, released in April 2026.

Name ensembles can be the same class of leftover: a reward signal or a fine-tune that made two names travel together, then a later release that tried to suppress them. Suppression is visible in the Claude table. It does not pull yesterday’s Zenodo deposit out of DataCite.

The same stickiness shows up outside universities. Elara Voss, GPT’s solo ghost, had no pre-LLM footprint and later appeared across dozens of Amazon listings. That is a commercial rhyme of the academic problem: a default name becomes a public identity because the model will not stop using it.

Zenodo Mints a DOI Before Anyone Reads the File

Zenodo is built for open deposits (data, code, preprints, posters) and it is free to use. Anyone can open an account. That design is the point of the repository, and it is also how 991 records can land in a month.

On 27 April 2026, during the same burst, Zenodo published a generative AI policy for depositors. It says all content must rest on research the authors actually did. AI tools cannot be listed as authors, because authorship means accountability and consent. Bulk uploads of AI text with no research basis are listed as unsuitable, as are fabricated results presented as experiments.

The policy draws on the COPE position on authorship and AI tools. Non-disclosure of AI use, on its own, is not grounds for removal. Fabricated results are. The March uploads were already in DataCite before that page went live. A rule that binds the next depositor does not automatically unwind a DOI that has already been minted.

WHAT THE IDENTIFIERS SHOW

  • 1,655 records: Ghost-authored Zenodo deposits with real DataCite DOIs and journals that do not exist.
  • 991 in March 2026: The peak month, against a 2025 baseline of one or two records a month in the same collection.
  • 23% pair rate: Elena Vasquez and Marcus Chen on Claude Sonnet 4 pair prompts, a search signature for Claude-made pages.
  • 35 papers, 0 citations: The Mei-Lin Zhang ResearchGate profile as of 15 May 2026, spanning disconnected fields.

The authors treat the names as a detection tool: probe the model, then search the web for the pair. That method works because the models are repetitive. It also explains why a later Claude release that drops Vasquez and Chen does not clean the databases. The fingerprint moves. The old DOI stays.

Identifiers Outlive the Models That Minted Them

Once a DataCite record exists, other systems can treat Elena Vasquez as a co-author with a citable output. Nothing in that chain asks whether she camped on Mount Nyiragongo in 2019 or whether Marcus Chen watched 11 eruptions on four continents. Those details were generated for a site that no longer exists.

A later model can also train on the deposit, then cite it, which is how a name prior becomes a loop. The March and April burst is a finished dump in the preprint. The identifier layer is not. Pulling a Zenodo record takes a takedown. Leaving it in place leaves a person who was never born inside the same metadata firehose that libraries, search tools, and citation trackers already trust.

Harry is the editor of SOMALI UPDATE, an independent title he owns and runs. Ten years in journalism, from reporter to editor, have settled into a set of verification habits he applies to every story. A quote is checked against the recording or transcript it came from. A statement attributed to an organisation is confirmed on that organisation's own channels before it is repeated. A figure is traced to the dataset or filing that first published it, and a photograph is checked for when and where it was actually taken. If any of those checks fails, the claim is left out or clearly marked as unconfirmed. Those habits cover the whole site, which reports news, business, technology, science and sports along with entertainment, lifestyle, travel, auto and gaming for readers around the world. Product claims in the technology, auto and gaming pages are tested in use where Harry can get his hands on the product. Corrections are published under a public policy and noted on the article. Readers who want to question a fact can write to support@somaliupdate.com.

Click to comment

Leave a Reply

Your email address will not be published. Required fields are marked *

Trending