Four recent papers expose separate failure points between retrieval and a trustworthy AI citation. SearchEngineAnswer reproduced PageRecall’s public tier-0 checks and built a claim-level audit ledger.
A new GMMM preprint connects generated-answer inclusion, question volume, system use and notice probability, but its empirical demonstration is simulated.
A dated compatibility test of six Schema.org 30.1 commerce properties finds a lagging public validator, successful Google parsing and a nested Product risk.
Chrome added four experimental ad metrics to CrUX. We measured a fixed publisher panel and found why count, density, CPU and network weight need separate diagnoses.
A 30,000-output study finds retrieval drives most variation while Reddit Answers favors early, top-level and more formal comments in its selected evidence.
A 712-query audit found detector-classified synthetic pages among citations, but 27.1% of cited URLs were not analyzed and page accuracy was not tested.
A source-long-tail recomputation plus a separate LinkedIn engine-mix example show why aggregate AI citation rankings can hide most domains and platform dependence.