







Don’t put faith in detectors that are “unreliable and easily gamed,” says scholar.
GPT detectors are biased against non-native English writers
GPT detectors frequently misclassify non-native English writing as AI generated, raising concerns about fairness and robustness. Addressing the biases in these detectors is crucial to prevent the marginalization of non-native English speakers in evaluative and educational settings and to create a more equitable digital landscape.

GPT detectors are biased against non-native English writers
GPT detectors frequently misclassify non-native English writing as AI generated, raising concerns about fairness and robustness. Addressing the biases in these detectors is crucial to prevent the marginalization of non-native English speakers in evaluative and educational settings and to create a more equitable digital landscape.

Literature fans should welcome AI as a fellow wordsmith | Aeon Essays
Strong resistance to AI among writers is understandable. But it obscures what we share with the machines: language itself

Can AI Detectors Be Trusted?
Society has always trusted editors, teachers, peer reviewers, and readers to set acceptable standards for writing. In the past, the biggest challenge was detecting plagiarism and cracking down on other forms of cheating. Now, AI detectors are rewriting the script: they assign a score based on the likelihood that the text was created by a human or a large language model (LLM).

People are getting their news from AI – and it’s altering their views
Even when information is factually accurate, how it’s presented can introduce subtle biases. As large language models increasingly bring people the news, this bias is a looming problem.

People are getting their news from AI – and it’s altering their views
Even when information is factually accurate, how it’s presented can introduce subtle biases. As large language models increasingly bring people the news, this bias is a looming problem.

Covert Racism in AI: How Language Models Are Reinforcing Outdated Stereotypes | Stanford HAI
Despite advancements in AI, new research reveals that large language models continue to perpetuate harmful racial biases, particularly against speakers of African American English.

AI assistants can sway writers’ attitudes, even when they’re watching for bias | Cornell Chronicle
Cornell Tech researchers found that writers who used biased AI auto-suggestions saw their views gravitate toward the AI’s positions without their realizing it — even when they were made aware of the biased AI.
Artificial Writing and Automated Detection
Artificial intelligence (AI) tools are increasingly used for written deliverables. This has created demand for distinguishing human-generated text from AI-generated text at scale, e.g., ensuring assignments were completed by students, product reviews written by actual customers, etc. A decision-maker aiming to implement a detector in practice must consider two key statistics: the False Negative Rate (FNR), which corresponds to the proportion of AI-generated text that is falsely classified as human, and the False Positive Rate (FPR), which corresponds to the proportion of human-written text that is falsely classified as AI-generated. We evaluate three leading commercial detectors—Pangram, OriginalityAI, GPTZero—and an open-source one —RoBERTa—on their performance in minimizing these statistics using a large corpus spanning genres, lengths, and models. Commercial detectors outperform open-source, with Pangram achieving near-zero FNR and FPR rates that remain robust across models, threshold rules, ultra-short passages, "stubs" (≤ 50 words) and ’humanizer’ tools. A decision-maker may weight one type of error (Type I vs. Type II) as more important than the other. To account for such a preference, we introduce a framework where the decision-maker sets a policy cap—a detector-independent metric reflecting tolerance for false positives or negatives. We show that Pangram is the only tool to satisfy a strict cap (FPR ≤ 0.005) without sacrificing accuracy. This framework is especially relevant given the uncertainty surrounding how AI may be used at different stages of writing, where certain uses may be encouraged (e.g., grammar correction) but may be difficult to separate from other uses.

Back-to-basics approach can match or outperform AI in language analysis
A new study led by Dr Andrea Nini at The University of Manchester has found that a grammar-based approach to language analysis can match or outperform advanced AI systems in identifying who wrote a text. The method, called LambdaG, uses patterns in grammar and sentence construction rather than large-scale AI models, offering comparable accuracy ...

Technical Report on the Pangram AI-Generated Text Classifier
We present Pangram Text, a transformer-based neural network trained to distinguish text written by large language models from text written by humans. Pangram Text outperforms zero-shot methods such as DetectGPT as well as leading commercial AI detection tools with over 38 times lower error rates on a comprehensive benchmark comprised of 10 text domains (student writing, creative writing, scientific writing, books, encyclopedias, news, email, scientific papers, short-form Q&A) and 8 open- and closed-source large language models. We propose a training algorithm, hard negative mining with synthetic mirrors, that enables our classifier to achieve orders of magnitude lower false positive rates on high-data domains such as reviews. Finally, we show that Pangram Text is not biased against nonnative English speakers and generalizes to domains and models unseen during training.

How easy is it to spot AI writing?
With AI writing becoming increasingly commonplace, how can we spot it and what difference does it make?

trying to articulate what i hate about "AI writing"
dan
so it's like if we imagine a normal distribution but then we set super intense resonance at exactly the middle point. that's how it feels to be surrounded by AI writing every day. it's average *without* the normal distribution — just extreme dose of narrow average. this is what pushes my buttons