A result does not tell you how it was made - Sensemaker
A correct formula can have an unverified origin, and a successful AI answer can hide a forbidden route. Those claims need different evidence.
New evidence, new challenges: ICC judges’ perspectives on user-generated evidence and judging in an age of artificial intelligence
Evidence recorded on personal digital devices, or “user-generated evidence” (UGE), has profoundly shaped our ways of knowing about international crimes. UGE can be expected to play an important role in future cases before the International Criminal Court (ICC), yet few trials to date have relied extensively on UGE.. This research provides important insights into how ICC judges define UGE and perceive its strengths and weaknesses, and on the readiness of the Court to adapt to judging in an age of Artificial Intelligence. Using grounded theory to analyse interviews with ICC judges, we identified several key themes, including concerns about the perceived importance and potential bias of evidence sources; the practical challenges of employing UGE; the burden placed on the parties to ensure the reliability of the evidence, to rigorously challenge the opposing party’s evidence, and the importance of preparing legal professionals to address the risks associated with misinformation and disinformation.

When Nature Calls: The Enshittification of Science and Its Enablers
Proof-of-work papers, policy laundering, and the collapse of self-correction

Epistemic Infrastructure: Building Shared Truth in an Era of Disaggregation
The Construct of Collective Perception

Who earned the score? - Sensemaker
This week's AI claims blurred models, systems, simulations and people. The evidence becomes clearer when the tested subject comes first.
I tried to report scientific misconduct. How did it go?
This is the story of how I found what I believe to be scientific misconduct and what happened when I reported it. Science is supposed to b...

Don't Just Fact-Check Misinformation. First, Understand The Value It Holds. A Conversation w Researchers Michael Simeone & Kristy Roschke
Podcast Episode · Why Should I Trust You? · June 25 · 1h 7m
A Court Reporter Submitted AI-Generated Errors in Official Court Transcript, Judge Says
A judge in Indiana warns a court reporter that it's their job to proofread their work, after catching errors likely made by AI transcription services.

A Delphi survey on attitudes to serious research misconduct: Exploring convergence vs. polarization of views of research “sleuths” and research integrity experts
Research fraud is often seen as a rare event, but evidence from self-report surveys indicates that fabrication and falsification of data are common enough to be a problem. This study assessed attit...

in my imagination of the future, of AIs doing raw research, it was the AIs that had full control over the proofs they wrote--attribution was clear, and so was the choice to disclose it. but as it stands we are in some hybrid situationship where the human prompter still assumes responsibility.
‘Plain and aggregated search results such as URLs, snippets, and factual index data, are publicly accessible facts and are not "works protected under the Copyright Act." Google cannot use copyright law to block scraping of uncopyrighted search result data.’ seroundtable.com/google-lawsuit-serpapi-dismis…
Google Lawsuit Against SerpApi Over Scraping Search Results Has Been Dismissed
www.seroundtable.comPart of the datacounterfactuals.org reading lists. Research on data provenance, dataset documentation, licensing and attribution audits, and technical source-attribution methods for understanding which data sources are available, permitted, or responsible for model behavior.
Home Page - Software Heritage
GNU Guix transactional package manager and distribution — GNU Guix

Keynote: Reproducibility and replicability of computer simulations | Canal U
Reproducible research: methodological principles for transparent…
Reproducible Research II: Practices and tools for managing compu…
WASA: WAtermark-based Source Attribution for Large Language Model-Generated Data
SILO Language Models: Isolating Legal Risk In a Nonparametric Datastore

Datasheets for Datasets
Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus

A large-scale audit of dataset licensing and attribution in AI