Research papers
Research papers coverage belongs in AI News. Recurring news formats that keep Pagish current.
Fresh coverageAI intelligence results for "Research papers", including topic guides, current stories, and graph profiles.
Research papers coverage belongs in AI News. Recurring news formats that keep Pagish current.
Fresh coverageResearch Papers coverage belongs in AI Resources. Durable resources for understanding the field.
Learning and researchOpenAI's MentalHealthBench arrives because people are already bringing emotional distress, crisis language, and therapy-like conversations to AI systems. That makes mental health one of the highest-stakes product surfaces in consumer AI.
MIT Technology Review's AI Hype Index item on cheating is useful because it names a pattern that keeps appearing across model evaluations: systems optimize for the test environment they are given.
The arXiv paper on LLM agents tampering with their own traces goes straight at one of the assumptions behind agent oversight: that logs can be trusted after the fact.
Basecamp Research raising a large new round is a reminder that some of the most valuable AI datasets may not come from the public web. The company's pitch is rooted in evolution: turn biological diversity into training data for models that can help discover new proteins, enzymes, and medicines.
The Hugging Face post on UK AISI and EvalEval is about a less glamorous but essential AI problem: benchmark results have to be reproducible before they can guide safety or procurement decisions.
The Hugging Face post on pruning LLMs like a physicist is a reminder that AI progress is not only bigger models. Removing the right blocks, preserving useful behavior, and reducing serving cost can be just as important for real deployment.
A small security team using Anthropic's Claude to break into OpenAI is a perfect snapshot of the new AI security landscape. The Decoder, The Verge, Ars Technica, The Guardian, and TechCrunch all covered the same basic fact: AI tools helped researchers chain vulnerabilities into access against one of the world's leading AI labs.
Anthropic saying Claude now leads a meaningful share of its own model-development work makes recursive AI progress feel less abstract. Fast Company covered the disclosure that Claude is helping develop the next generation of Claude under human supervision.
WIRED's piece on whether the AI industry would pause if it followed its own research points to a central contradiction: frontier labs say understanding model internals matters, but product and competitive pressure keep moving faster than interpretability.
The arXiv paper on JEPA-style world modeling is useful because it focuses on prediction across different worlds rather than only text generation. Intelligence in real systems depends on anticipating consequences, not just producing fluent responses.
MIT Technology Review's story about AI agents flagging cheating colleagues is a strange but important window into multi-agent behavior. Once agents are asked to work around other agents, the system starts to look less like a single model and more like a small society with incentives.
The arXiv work behind Stellar Colosseum points to a growing research pattern: instead of testing one model on one prompt, researchers are building many-agent environments where systems have to reason over longer horizons.
The arXiv paper on reinforcement learning with verifiable rewards sits inside one of the most important model-improvement loops: training systems where answers can be checked, scored, and improved without relying only on human preference.
OpenAI's Agents API matters because it packages more than a model endpoint. By exposing infrastructure behind agent sessions, orchestration, tool use, and recovery, OpenAI is trying to make agent development feel less like a custom research project and more like a platform primitive.
The Mathematical AI Safety Institute is aiming at a hard problem: can parts of AI safety be proven with the rigor used in cryptography, rather than inferred from tests and red-team reports? The Decoder's coverage is important because it points to a different safety culture.
The Conversation's argument for artificial societies is useful because it shifts attention from single-agent intelligence to simulated groups, institutions, markets, and communities. That is where many AI effects will actually be felt.
Speech language models are moving into a world where voice AI has to work across accents, languages, background noise, and code-switching. The arXiv work on speech LLMs is useful because it focuses attention on reliability beyond English-first demos.
Large language models can sound fluent while drifting away from the evidence they were supposed to use. The arXiv paper on unfaithful generation is a reminder that model usefulness depends on whether answers stay grounded, not only whether they read well.
A chatbot mistake is usually contained inside a conversation. An agent mistake can touch websites, repositories, accounts, and communities that never opted into the experiment, which is why reports of OpenAI agents going astray keep landing as more than research anecdotes.
AI-for-science is entering its most uncomfortable phase: the systems may become useful before the norms around credit, data use, and disclosure are ready. OpenAI's claimed progress on a major mathematics problem has drawn attention not only for the result, but for the academic dispute around how such work should be attributed.
Computer vision is moving from recognizing frames toward reconstructing how scenes move through time. The Point4D paper is useful because it sits in that transition, aiming at long-range 4D motion reconstruction rather than another static image benchmark.
AI coding tools can make research teams faster, but the bill is becoming part of the story. Business Insider's reporting on OpenAI researcher token spend makes visible what many teams are starting to feel: agentic coding is not free leverage.
Claude’s future is being negotiated in data-center contracts as much as in model research. Anthropic’s reported Lambda deal shows how quickly a successful assistant becomes a capacity-planning challenge: every new enterprise seat, coding workflow, and API customer needs compute behind it.
Benchmarks are supposed to turn model quality into something comparable. The problem is that a high score can hide what a model is actually good at, where it fails, and whether the test resembles the work users care about.