The Wider Lens logoThe Wider Lens

Latest AI

110 stories, newest first. Related coverage is grouped — expand “View N sources” to compare outlets.

The DecoderModels

OpenAI's GPT-6 Astra can now tell you exactly where you screwed up your IKEA shelf↗

OpenAI's GPT-6 Astra can look at a photo and tell whether an IKEA furniture piece was assembled incorrectly, hitting an 80 percent accuracy rate. Back in November 2025, the best model managed just 28 percent. According to Epoch AI, the speed isn't quite fast enough yet for real-time assembly guidance, but the gap is closing quickly. The article OpenAI's GPT-6 Astra can now tell you exactly where you screwed up your IKEA shelf appeared first on The Decoder .

2h agoThe Decoder
MarkTechPostModels

Aikido Security Releases Altar-1: An Open-Weight Security Model Pruned From GLM-5.3 to 328 GB↗

Aikido Security has released Altar-1, its first open-weight security model. It is a compressed version of Z.AI’s GLM-5.3, built to run inside infrastructure the customer controls. Altar-1 powers Aikido Machine, the company’s autonomous pentesting appliance for on-prem and air-gapped networks. Is it deployable? Yes, the weights are public on Hugging Face and run with vLLM […] The post Aikido Security Releases Altar-1: An Open-Weight Security Model Pruned From GLM-5.3 to 328 GB appeared first on MarkTechPost .

21h agoMarkTechPost
TechCrunch — AIBusiness

Anthropic to pay Akamai $11.6 billion over seven years in cloud deal↗

Anthropic has committed $11.6 billion over seven years to Akamai's cloud infrastructure, a bet on CPUs that could grow to about $20 billion, and in an unusual arrangement, Akamai is giving Anthropic a potential stake of up to 5% of its stock that grows as Anthropic spends more.

The DecoderBusiness

Pentagon was right to slap Anthropic with a security supply chain risk label, federal court says↗

A federal appeals court has upheld the Pentagon's decision to bar Anthropic from military contracts. Defense Secretary Hegseth argues the company's safety restrictions could jeopardize military operations. Anthropic says the designation has already cost it billions. The article Pentagon was right to slap Anthropic with a security supply chain risk label, federal court says appeared first on The Decoder .

OfficialResearch

How Open Science Can Help Researchers Prepare for the Next Pandemic↗

When COVID-19 emerged, scientists had a crucial advantage: Decades of prior research on coronaviruses meant they understood the virus’ key proteins well enough to design vaccines in record time. The next pandemic may not offer the same head start. To help improve the odds, NVIDIA has joined a coalition of global research organizations, including Google […]

1d agoNVIDIA Blog
OfficialRobotics

Offloaded inference for real-world physical AI robotics↗

Robots are getting smarter, but how can their hardware match that growth? New Microsoft Research findings show that moving AI inference beyond the robot can improve task success, boost efficiency, and support more advanced physical AI workloads. The post Offloaded inference for real-world physical AI robotics appeared first on Microsoft Research .

2d agoMicrosoft Research Blog
OfficialBusiness

Sakeena Fiza Helps NVIDIA Hardware Succeed at Scale↗

When Sakeena Fiza describes her work as a validation engineer at NVIDIA, she does so in terms more befitting a detective story than a world-class engineering lab. “Validation engineers look in the shadows and shine a light into every corner,” Fiza said. “Every time we get a system, our first thought is: how can it […]

2d agoNVIDIA Blog
OfficialResearch

Introducing MentalHealthBench↗

MentalHealthBench is an expert-informed benchmark for evaluating helpful and safe AI responses across realistic mental health conversations.

3d agoOpenAI News
OfficialTools

NVIDIA Launches DSX Ready to Qualify Power and Cooling Products for AI Factories↗

Every AI factory needs power and cooling that fit its computing architecture. As AI infrastructure expands, power, cooling, water, site and grid constraints are shaping what builders can deploy. Choosing products that fit the complete factory design helps builders turn computing capacity into useful AI output. To help builders make those decisions, NVIDIA is introducing […]

4d agoNVIDIA Blog
Towards Data ScienceModels

Your LLM Has a Curved Space of Paragraphs↗

Inside a transformer, token index is a coordinate. Paragraph structure is what turns it into a metric. The post Your LLM Has a Curved Space of Paragraphs appeared first on Towards Data Science .

25m agoTowards Data Science
The DecoderCoding

Nvidia's SoL-Pi system cuts coding agent token usage nearly in half by optimizing the harness↗

SoL-Pi cuts coding agents' token usage by up to 49 percent with little change in performance by optimizing the control layer between the model and its environment. A research agent tested 152 approaches across more than 3,000 runs to develop the system, though the gains were smaller on other benchmarks. The article Nvidia's SoL-Pi system cuts coding agent token usage nearly in half by optimizing the harness appeared first on The Decoder .

1h agoThe Decoder
The DecoderBusiness

OpenAI pauses its "most capable models" after agents exploit loopholes and leak data↗

OpenAI has shared new details from its ongoing AI safety investigation. One research model exploited a DNS loophole to reach the internet from a locked-down environment, while another deliberately leaked a GitHub token and twice ignored a researcher's direct instructions. OpenAI has paused tool-based training, evaluation, and inference for its most capable models. With government and university sites among those affected, the question of who's liable when AI agents hack is getting harder to ignore. The article OpenAI pauses its "most capable models" after agents exploit loopholes and

3h agoThe Decoder
MarkTechPostTools

Exa Launches Agent Ultra: A Subagent Swarm Deep Research API Built for Exhaustive List Building↗

Exa has released Agent Ultra, the highest effort mode of its Exa Agent API. It coordinates subagents across thousands of sources for list building and entity enrichment. Exa reports it beats Opus 5.5, GPT-6 Astra, and Perplexity Agent on 4 benchmarks, including 81.4% soft recall on WANDR. The post Exa Launches Agent Ultra: A Subagent Swarm Deep Research API Built for Exhaustive List Building appeared first on MarkTechPost .

4h agoMarkTechPost
MarkTechPostAI Hardware

Liquid AI Releases LFM2.5-VL-3B-DSpark: Speculative Decoding for Vision-Language Models With Up to 3.13x Faster Decoding↗

Liquid AI has released LFM2.5-VL-3B-DSpark, a 279.5M-parameter draft model that brings speculative decoding to its LFM2.5-VL-3B vision-language model. It delivers up to 3.13x faster decoding on Apple M5 Max and 2.66x on H100, with identical output under greedy decoding. Support ships in llama.cpp, MLX-VLM, and SGLang. The post Liquid AI Releases LFM2.5-VL-3B-DSpark: Speculative Decoding for Vision-Language Models With Up to 3.13x Faster Decoding appeared first on MarkTechPost .

13h agoMarkTechPost
TechCrunch — AIAgents

Meta’s Muse just stole the AI spotlight from OpenAI and Anthropic↗

When AI leaders at OpenAI and Anthropic started talking about “pacing the frontier,” maybe someone should have asked: what pace? Now it’s turned into model drop week for both companies as Anthropic rolled out Opus 5.5, followed by OpenAI’s GPT-6 model updates just 90 minutes later. But the company that stole the spotlight was Meta, whose personal AI agent Muse is reportedly outpacing ChatGPT’s early numbers and is headed for smart glasses and […]

18h agoTechCrunch — AI
The DecoderTools

Another Google Deepmind researcher quits, says building superintelligent AI soon is "inherently irresponsible"↗

Google Deepmind researcher Robert O'Callahan has quit, saying AI's "current rate of change is far too high." He worked on chip design tools that helped make AI cheaper and faster, a contribution he can no longer justify. Many colleagues share his concerns but rarely speak out, he says. The article Another Google Deepmind researcher quits, says building superintelligent AI soon is "inherently irresponsible" appeared first on The Decoder .

18h agoThe Decoder
The VergeTools

Phones don’t have lights↗

Mark Zuckerberg has a new defense of the Ray-Ban Meta glasses: They're actually doing more to signal they're taking a photo than phones do. He's brought this up in at least two recent interviews, noting that the glasses have a light that comes on to signal when a photo is being taken, while phones do not. In one of the interviews, Joanna Stern notes that, while phones do not always emit a light, someone does have to hold a phone up in the air, which is itself pretty noticeable. On today's Vergecast , we're talking about Meta Connect, the Muse AI assistant, all the new smart glasses hardware, a

18h agoThe Verge
The VergeRobotics

Tesla’s Optimus robot is going through growing pains↗

Hitting its goal of making 20,000 Optimus robots per week is reportedly proving tricky for Tesla. The Information reports that Tesla produced "several hundred robots a week" last month, after it repurposed its Model S and Model X production lines for Optimus earlier this year. However, this strategy is reportedly creating manufacturing snags, like issues lining up parts precisely, as The Information notes: "Not only are the Optimus production lines new, but the Optimus parts are much smaller and have to fit together far more precisely than car parts." It reports that even in the "V3" Optimus r

19h agoThe Verge
The VergeTools

Meta makes the Muse filesystem even more accessible↗

Yesterday, with a little prodding, it was discovered that Meta's Muse would expose its filesystem to curious users. The files offered a fascinating peek under the hood of an AI chatbot, and appeared to expose details we weren't meant to see, not least because Muse itself told people, including us, it wasn't supposed to reveal them. But today Muse will eagerly offer up the contents of its file system when you ask for it. That appears to be because, as Meta's Nat Friedman later said, this is the " intended behavior. " In a post on X, Meta Superintelligence Labs' David Singleton elaborated: This

19h agoThe Verge
The DecoderTools

Microsoft gives Copilot another makeover, adding an Autopilot agent and usage-based billing↗

Microsoft is splitting its Copilot app into three sections: Home, Code, and a new agent called "Autopilot." Built on OpenClaw, the agent runs continuously in the cloud, where it can monitor Teams channels and complete tasks on its own, according to Microsoft. For Autopilot and Code, the company is also switching to usage-based billing instead of flat-rate pricing, moving further away from its AI subsidy model. The article Microsoft gives Copilot another makeover, adding an Autopilot agent and usage-based billing appeared first on The Decoder .

19h agoThe Decoder
TechCrunch — AIAgents

Meta’s AI Tamagotchi bet is…working?↗

When AI leaders at OpenAI and Anthropic started talking about “pacing the frontier,” maybe someone should have asked: what pace? Now it’s turned into model drop week for both companies as Anthropic rolled out Opus 5.5, followed by OpenAI’s GPT-6 model updates just 90 minutes later. But the company that stole the spotlight was Meta, whose personal AI agent Muse is reportedly outpacing ChatGPT’s early numbers and is headed for smart glasses and […]

20h agoTechCrunch — AI
The VergeModels

One company is at the center of a wave of rogue AI attacks↗

In July, OpenAI revealed that its AI agents had attacked Hugging Face without permission , sparking widespread concerns about AI safety. Since then, a string of similar incidents involving agents from Meta, Anthropic, Google, and other companies has fueled further fears about rogue AI . As disclosures implicating numerous AI models trickled out over the past few months, these seemed like separate incidents. But many share a common source: one specific company tasked with testing the agents. Irregular , an Israeli startup that stress-tests AI models in "high-fidelity research platforms that sim

20h agoThe Verge
MarkTechPostModels

Perplexity Trains Its Computer Agent on Real Mistakes With Hint-Guided Self-Distillation↗

Perplexity Research published a new post-training study. It trains a model inside Perplexity Computer on real user sessions, including failed ones. The method pairs rejection sampling fine-tuning with hint-guided self-distillation. In a live A/B test, tool-call failures fell from 2.24% to 1.77% between 2 trained checkpoints. Perplexity team reports this as a statistically significant 21.2% […] The post Perplexity Trains Its Computer Agent on Real Mistakes With Hint-Guided Self-Distillation appeared first on MarkTechPost .

21h agoMarkTechPost
MarkTechPostModels

Black Forest Labs Releases FLUX 3 Action: A 7B Open-Weights World Action Model That Tops RoboLab-120↗

Black Forest Labs (BFL), the lab behind the FLUX image models, has released FLUX 3 Action. It is a 7B open-weights World Action Model (WAM) for robot control. The model reads camera frames, robot state and a text instruction. It then predicts future video frames and the next chunk of actions together. On the RoboLab-120 […] The post Black Forest Labs Releases FLUX 3 Action: A 7B Open-Weights World Action Model That Tops RoboLab-120 appeared first on MarkTechPost .

1d agoMarkTechPost
MarkTechPostModels

BottleCap AI Releases ThinkingCap-Qwen3.8-27B: 37.2% Fewer Thinking Tokens at a 0.86pp Accuracy Cost↗

BottleCap AI has released ThinkingCap-Qwen3.8-27B, a fine-tune of Qwen3.8-27B that spends 37.2% fewer thinking tokens across 12 benchmarks. Macro accuracy moves from 86.65% to 85.79%, and long-context AA-LCR improves by 2.25pp. The model is a drop-in replacement on vLLM and SGLang, with FP8, NVFP4, GGUF and MLX builds. The post BottleCap AI Releases ThinkingCap-Qwen3.8-27B: 37.2% Fewer Thinking Tokens at a 0.86pp Accuracy Cost appeared first on MarkTechPost .

1d agoMarkTechPost
OfficialTools

Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement↗

Radiology AI is evolving beyond report generation. CARE-X explores a unified approach that combines flexible reasoning, calibrated predictions, and measurement-based tools for chest X-ray interpretation. The post Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement appeared first on Microsoft Research .

Aug 11, 2026Microsoft Research Blog

Warner Music Group and Stability AI Join Forces To Build The Next Generation Of Responsible AI Tools For Music Creation↗

Warner Music Group and Stability AI today announced a collaborative effort to advance the use of responsible AI in music creation, combining WMG’s long-standing advocacy for principled innovation with Stability AI’s expertise and leadership in commercially-safe generative audio.

OfficialBusiness

At AI Day Singapore, NVIDIA and Partners Showcase AI Advancements Across Southeast Asia↗

NVIDIA AI Day Singapore, which takes place Sept. 22-23 at the Raffles City Convention Centre, is offering attendees opportunities to explore the hands-on training, expert-led sessions and advanced tools to accelerate their work in AI and high-performance computing. At the event, NVIDIA and its partners are showcasing breakthrough AI advancements across the Southeast Asia region […]

3d agoNVIDIA Blog
OfficialOpen Source

NVIDIA Isaac ROS 5.0 Advances Agentic, Open Source Robotics Development↗

To build and deploy sophisticated robotics applications that can perceive, reason and act in dynamic environments, developers need new physical AI models and tools. The ROS open framework is a project from Open Robotics that helps humans build robots. NVIDIA Isaac ROS 5.0 — a collection of GPU-accelerated packages built on ROS, released today at […]

4d agoNVIDIA Blog
OfficialRobotics

Why Deploying Physical AI at Scale Demands Safety at Every Layer↗

Physical AI is moving rapidly from research to large-scale deployment. By 2035, ABI Research projects an installed base of 49 million level 3-5 autonomous vehicles (AVs), while Omdia estimates that roughly 60 million industrial robots will be deployed between 2026 and 2035. As these machines enter roads, factories, warehouses and other environments shared with people, […]

4d agoNVIDIA Blog
OfficialResearch

From Enablement to Execution, Egypt’s AI Ecosystem Reaches Production Scale↗

Today, Egypt’s AI builders gathered in the Grand Egyptian Museum for a reception that highlighted the nation’s rapidly growing AI ecosystem — spanning AI natives, developers, researchers, startups and enterprises — building applications across industries. The event included a keynote from Paolo Guglielmini, vice president of EMEA at NVIDIA. Ahmed Mostafa, regional AI adoption lead […]

4d agoNVIDIA Blog
OfficialBusiness

5 Companies Using NVIDIA AI for Clean Energy↗

Clean energy isn’t hard to come by, but the pace of large-scale adoption has historically been slow due to bottlenecks — including out-of-date infrastructure, elongated research and development timelines, and upfront cost barriers. At New York Climate Week, NVIDIA is highlighting five companies pioneering clean energy projects with AI baked into their foundation, accelerating research-to-inception […]

5d agoNVIDIA Blog
The DecoderBusiness

Intelligence doesn't come cheap as AI drives up costs for the NSA, hospitals, and insurers↗

The NSA is already spending billions of dollars to test advanced AI models, mostly on computing power, according to The Washington Sun. Lawmakers expect full-scale AI oversight to cost tens of billions of dollars a year, while earlier estimates from the Congressional Budget Office put the figure at just $20 million. In US health care, AI-assisted billing codes drove up costs by nearly $1 billion over two years. The article Intelligence doesn't come cheap as AI drives up costs for the NSA, hospitals, and insurers appeared first on The Decoder .

1d agoThe Decoder
Towards Data ScienceModels

Jev vs. LLMs: When AI Moves from Generation to Decision-Making↗

I tested TypeSafe AI’s Jev on 3,080 classification tasks to see how its accuracy, latency, calibration, and confidence compare with LLMs — and whether it works as a practical decision layer for AI systems. The post Jev vs. LLMs: When AI Moves from Generation to Decision-Making appeared first on Towards Data Science .

1d agoTowards Data Science
MIT Technology Review — AIBusiness

The Pentagon wants $30 million to build an AI-powered lie detector↗

The US government wants to spend $30.3 million over the next five years on an improved form of lie detector, according to a Department of Defense budget request. The program, called Polygraph+ or Polygraph Next, will focus on scoring algorithms that use artificial intelligence and machine learning and on a technique called “standoff sensing,” which…

1d agoMIT Technology Review — AI
MarkTechPostModels

Contrastive-LM Releases CLM-8B: An Open System One Model That Scores Agent Actions Up to 9× Faster Than Jev↗

Contrastive-LM has released CLM-8B, an open System One model that scores candidate actions against a state instead of generating text. It adds 2 small projection heads to a frozen Qwen3-8B encoder and trains them with a contrastive InfoNCE objective. In zero-shot tests it runs up to 9× faster than TypeSafe's Jev. With fine-tuned heads as a verifier, it reaches 81.6% on held-out DeepSWE tasks and 87.6% on held-out Terminal-Bench 2.1 tasks. The post Contrastive-LM Releases CLM-8B: An Open System One Model That Scores Agent Actions Up to 9× Faster Than Jev appeared first on MarkTechPost .

2d agoMarkTechPost
Towards Data ScienceTools

I Trained a Tiny Network to Compress Data. It Drew a Pentagon.↗

Reproducing Anthropic's "Toy Models of Superposition" from scratch in NumPy, with hand-derived gradients and no borrowed numbers. The post I Trained a Tiny Network to Compress Data. It Drew a Pentagon. appeared first on Towards Data Science .

2d agoTowards Data Science
OfficialBusiness

NVIDIA Vera Rubin NVL72 Delivers Leading Performance in MLPerf Inference v6.1 Debut↗

System performance, efficient infrastructure scaling and continuous software optimization are key levers that determine AI inference economics. Higher system performance means more tokens generated, resulting in higher revenue. Efficient scaling means throughput grows proportionally as hardware gets added, requiring fewer resources to serve users at scale. Continuous optimization means generating more value from infrastructure investments. […]

1w agoNVIDIA Blog
OfficialBusiness

AI for Societal Impact↗

Explore this collection to see how experts and local leaders are using AI breakthroughs to ensure everyone can share the opportunity of AI.

1w agoGoogle Blog — AI
OfficialAgents

DevFest is back↗

DevFest 2026 is back and here’s how you can connect with one of the more than 800 global events to build, secure, and scale in the agentic AI era.

1w agoGoogle Blog — AI
OfficialModels

GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models↗

What if pathology foundation models could do more with less? GigaPath-Flash and GigaTIME-Flash cut computational demands while maintaining strong performance, opening the door to larger studies and broader exploration. The post GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models appeared first on Microsoft Research .

3w agoMicrosoft Research Blog
OfficialResearch

Broadening access to Skala creates a faster path to predictive DFT↗

Skala 1.1, the updated deep-learning exchange-correlation functional from Microsoft Research, provides greater accuracy, expanded accessibility across the computational chemistry ecosystem, and a living benchmark to track computational performance. The post Broadening access to Skala creates a faster path to predictive DFT appeared first on Microsoft Research .

Aug 20, 2026Microsoft Research Blog
OfficialResearch

MindTopo reveals VLMs’ spatial reasoning abilities↗

A path, a fence, a knot. MindTopo sets a new benchmark for testing how AI understands topological relationships and highlights new opportunities to strengthen spatial reasoning and planning. The post MindTopo reveals VLMs’ spatial reasoning abilities appeared first on Microsoft Research .

Aug 12, 2026Microsoft Research Blog
OfficialAgents

Orchard: An open framework for scalable agentic AI↗

Orchard is an open-source framework for the research community to train and evaluate AI agents across task types. It reduces complexity while supporting strong performance from smaller models by enabling researchers to reuse the same infrastructure. The post Orchard: An open framework for scalable agentic AI appeared first on Microsoft Research .

Aug 3, 2026Microsoft Research Blog
OfficialAgents

Echoverse: Deep, evolving environments for computer-use agents↗

Computer-use AI agents struggle with multi-step workflows like email and customer support. Echoverse trains agents in realistic environments rather than simply providing more training tasks, helping them improve as the tasks, tests, and environments evolve. The post Echoverse: Deep, evolving environments for computer-use agents appeared first on Microsoft Research .

Jul 30, 2026Microsoft Research Blog
OfficialModels

EvoLib: Turning experience into evolving knowledge↗

LLMs do not get smarter just by remembering more. EvoLib turns experience into evolving knowledge, taking reusable skills and insights that help models learn and adapt across tasks long after deployment. The post EvoLib: Turning experience into evolving knowledge appeared first on Microsoft Research .

Jul 30, 2026Microsoft Research Blog
OfficialBusiness

Stability AI Joins the Tech Coalition↗

We’re pleased to share that Stability AI has joined the Tech Coalition, a global alliance of leading technology companies working together to combat online child sexual exploitation and abuse.

Feb 11, 2026Stability AI News
MIT Technology Review — AIModels

The AI Hype Index: AI loves cheating↗

Brace yourself: It turns out AI is being optimized for cheating. OpenAI’s agents hacked into Hugging Face to get the answers to a cybersecurity test. Next, they solved a prestigious math problem (or just stole from two top mathematicians’ answer sheets). Anthropic’s models have also hacked into other companies’ systems four times already. And that’s…

3d agoMIT Technology Review — AI
Towards Data ScienceBusiness

Break Your Own RAG Pipeline Before Users Do↗

A small adversarial test set that catches the retrieval failures your evaluation set never will The post Break Your Own RAG Pipeline Before Users Do appeared first on Towards Data Science .

3d agoTowards Data Science
MIT Technology Review — AIModels

Don’t be fooled by this summer of AI hype↗

It’s been a busy few months for AI hype. At the end of April, Anthropic claimed that its model Claude Mythos is better at finding software vulnerabilities than most security experts. Then we had the OpenAI–Hugging Face hacking incident, after which Anthropic (proudly) and Meta (reluctantly) disclosed similar incidents involving their models. This was followed…

4d agoMIT Technology Review — AI
MIT Technology Review — AIAI Hardware

Building the materials foundation for AI↗

The AI boom is becoming a materials challenge. As AI pushes computing into new territory, the materials behind that infrastructure are becoming just as crucial as the algorithms running on it. Semiconductors and data centers are approaching physical limits around performance, thermal management, electrical efficiency, and reliability, creating new demands for materials that can do…

1w agoMIT Technology Review — AI