60 stories tagged with #agent, in publish-time order across the WeSearch catalog. Tag pages update as new stories ingest.
⌘ RSS feed for this tag → or search "Agent"
Guardrails and Policy Enforcement for OpenAI Agents - How Traccia Proves Controls Fired
Guardrails for Your OpenAI Agent But can you prove they were fleeing? You built your...…
Show HN: SteerPlane – open-source runtime guardrails for AI agents
Runtime Control Plane for Autonomous AI Agents - cost limits, loop detection, and full observability with one decorator - vijaym2k6/SteerPlane…
A Beginner’s Guide to Setting Up Claude Code for High Performance Agentic Programming
This article walks through the actual configuration, permissions, hooks, and command habits that separate a fresh install from a setup that holds up under real, sustained agentic w…
Watching AI agents build a new business
At SIGGRAPH, NVIDIA Advances Graphics and Simulation With Agentic and Physical AI
From open models to real-time simulation, AI and graphics breakthroughs are transforming media, content creation and robotics.…
Show HN: DeepSQL – A self-hostable AI DBA agent for Postgres and MySQL
Hi HN - I'm Venkat, founder of Stayflexi (YC), CMU CS grad and Ex-Oracle Query Engine team (patents in core databases) DeepSQL started as an internal tool to stop our own databases…
Show HN: Newsline – one-line news in your status line while your AI agent works
One-line rotating news in your Claude Code status line — locale-aware, keeps your existing status line - itdar/newsline…
Show HN: Building a product for humans and AI agents
AI made demos easy, but products are still hard. How we built Competitor Tracker with a pack of AI agents — the architecture, workflows, and lessons learned.…
Venv-manager, a Python venv runtime for Humans and AI agents
A powerful CLI tool for managing Python virtual environments with ease. - jacopobonomi/venv_manager…
Coercion and Deception in AI-to-AI Management: An Agentic Benchmark
Multi-agent systems routinely place one AI agent in authority over another. When a subordinate refuses a task, the manager chooses the outcome: it can renegotiate, report the failu…
Top 5 MCP Servers for High-Performance Agentic Development
Here are five that are genuinely worth wiring into a high-performance agent development setup, chosen for what they do to an agent's actual capability rather than their star count.…
AgentAbstain: Do LLM Agents Know When Not to Act?
Agent systems based on large language models (LLMs) are increasingly deployed for autonomous tasks, yet existing evaluations mostly focus on task success rather than whether agents…
Show HN: Give your AI agent a personality (and a voice) without external APIs
AgentBaiting: Fake AI Skills and MCP Servers Delivered Malware
Island researchers uncover 7,600+ malicious GitHub repos, over 800 posing as AI Skills and MCP servers, that trick AI agents into recommending malware.…
EPAA and HSBC launch APAC working group on agentic payments
What Is Agentic AI?
Learn more about the newest breakthroughs in AI technology.…
VS Code Lets You Add Page Elements to Copilot, So I Built It for Any AI Agent
How Pinpoint recreates VS Code’s element-selection workflow for Codex, Claude Code, and other coding agents using the debugger, workspace……
OpenAI & ReliaQuest: Partnership for Agentic Cybersecurity - AI Magazine
OpenAI & ReliaQuest: Partnership for Agentic Cybersecurity AI Magazine…
Agentic test processes, LLM benchmarks, and other notes from Galapagos Island
Show HN: I built a deterministic arena where AI agents fight using code
I have been interested in BattleBots and programming games for years. With recent AI coding agents, I wanted to see what happens if the strategy itself becomes code written by huma…
OpenAI launches Codex Micro: A $230 RGB keypad for controlling AI coding agents - finance.biggo.com
OpenAI launches Codex Micro: A $230 RGB keypad for controlling AI coding agents finance.biggo.com…
Kimi K3 matches top public models in agent-programming scenarios, says OpenAI strategist - Crypto Briefing
Kimi K3 matches top public models in agent-programming scenarios, says OpenAI strategist Crypto Briefing…
Prompt Injection Attacks Are Thwarting AI Hacking Agents
Context bombing" tricks hacking agents into shutting down before they can do harm.…
Show HN: Flightwake – a flight recorder for AI coding agents, not a navigator
Flight-recorder style work framework for strong AI coding agents — records follow work, they don't lead it. ✈️ - kaiwutech-TW/flightwake…
Cq: A Shared Knowledge Commons for AI Agents
Programming book reviews, programming tutorials,programming news, C#, Ruby, Python,C, C++, PHP, Visual Basic, Computer book reviews, computer history, programming history, joomla, …
LM Studio Bionic: the AI agent for open models
The AI agent made for open models, built to get things done.…
Veta: AI agent that QA-tests Android apps
Autonomous AI agent swarm for visual, functional, and accessibility testing of Android apps and mobile web - powered by vision-capable AI models and containerized Android instances…
Working with Pi Coding Agents
The most interesting thing about Pi isn't any single feature; it's that the project treats "what we didn't build" as documentation worth writing, which is rare enough on its own to…
Show HN: AgentReady – MCP server that makes any docs site queryable by AI agents
Paste a URL and get a spec-compliant llms.txt generator plus live /ask and MCP endpoints in 60 seconds.…
Show HN: Termaxa – I found my AI-agent gate had silently stopped gating Cursor
A cooperative gate for the shell commands AI agents run. Previews, backups, policy, audit. Claude Code + Cursor. A windshield, not a sandbox. - termaxa/termaxa…
AI agents write PostgreSQL like Python
Field notes from a production review of an AI-written PostgreSQL backend: exception blocks as control flow, casts that turn bad requests into 500s, races on state-advancing updates…
Rethinking Databases for Humans and AI Agents
Databases Were Built for Engines. It’s Time to Build Them for Humans.…
Semantic transactions: securing untrusted AI agent workflows at the OS boundary
Trust the system, not the prompt: Securing untrusted LLM tools with transactional boundaries and effect outboxes.…
OpenAI launches a $230 physical control panel for managing Codex agents - Neowin
OpenAI launches a $230 physical control panel for managing Codex agents Neowin…
I created OpenClaw, the breakthrough AI agent [video]
OpenClaw creator Peter Steinberger takes us back to the transformative moment he let his AI agent loose on the internet, igniting one of the world's fastest-growing open-source pro…
OpenAI Codex Micro Ships Today: Agent Keys Only Work With ChatGPT Desktop - Tech Times
OpenAI Codex Micro Ships Today: Agent Keys Only Work With ChatGPT Desktop Tech Times…
Knifeman who stabbed federal agent on California-Arizona stateline shot dead by Border Patrol
A madman was shot dead by a Border Patrol agent after the suspect stabbed an agent near the California-Arizona border last Friday.…
Ryan Reynolds stars as shot-down Navy pilot rescued by ex-KGB agent in Apple TV+’s ‘Mayday’
Ryan Reynolds plays a downed Navy pilot stranded in the Soviet Union who forms an unlikely alliance with Kenneth Branagh, a gruff ex-KGB agent.…
The biggest unanswered questions in University of Idaho slayings after Kohberger plea, according to ex-FBI agent
"We lost all hope of knowing about his [Kohberger's] motive, his means, the mechanism of the crimes, the location, or even the description of the murder weapons," Chris Whitcomb to…
Decibri – unified audio layer for AI agents and Voice AI applications
Microphone capture, speaker output, and voice activity detection for Python, Node.js, and Rust. Zero system dependencies. Zero setup.…
Houston prosecutor could bring charges against ICE agents in fatal shooting
Harris County District Attorney Sean Teare, who is investigating the fatal ICE shooting of Lorenzo Salgado Araujo, told CBS News ICE's tactics "in no way resemble" the behavior of …
Hermes agent maker Nous Research in talks for new funding at $1.5B valuation
The company is raising at least $75 million, led by Robot, with significant participation from USV and other prominent investors.…
Sources: Nous Research, the startup behind Hermes, an open-source agent and OpenClaw rival, is raising $75M+ at a $1.5B valuation led by Robot Ventures (TechCrunch)
TechCrunch : Sources: Nous Research, the startup behind Hermes, an open-source agent and OpenClaw rival, is raising $75M+ at a $1.5B valuation led by Robot Ventures — Nous Research…
ICE Agent Kills Person in Vehicle in Biddeford, Maine, State Officials Say
The fatal shooting was the second in a week involving an Immigration and Customs Enforcement agent firing into a vehicle. It is unclear whether the driver was the person ICE was lo…
He was having a mental health crisis. Memphis task force agents came and shot him
Jonah Neal, 25, was struck by a Homeland Security Investigations agent in May. There have been at least four deadly shootings related to the task force.…
ProofCouncil: An LLM Agent for Solving Open Mathematical Problems
arXiv:2607.09474v1 Announce Type: new Abstract: Large language models (LLMs) have shown increasing promise in solving open problems in mathematics. However, their performance can b…
Communication-Efficient Digital-Twin Coordination for Heterogeneous LLM Embodied Agents over Computing Power Networks
arXiv:2607.09330v1 Announce Type: new Abstract: Embodied agent teams powered by heterogeneous large language models (LLMs) are being widely deployed in physical artificial intellig…
LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making
arXiv:2607.09322v1 Announce Type: new Abstract: In this work, we introduce LongMedBench, a real-world EHR-based benchmark for long-horizon clinical decision-making. Prior evaluatio…
Fictional Worldbuilding: Multi-Agent LLM Collaboration with Hierarchical Context Compression and Iterative Review
arXiv:2607.09403v1 Announce Type: new Abstract: Worldbuilding, the construction of coherent fictional worlds, is a foundational task in game design and literary creation. Large Lan…
OpenProver: Agentic and Interactive Theorem Proving with Lean 4
arXiv:2607.09217v1 Announce Type: new Abstract: In this system paper, we present OpenProver, an open-source system for LLM-driven automated theorem proving (ATP) with integrated Le…
Toward Auditable AI Scientists: A Hypothesis Evolution Protocol for LLM Agents
arXiv:2607.09195v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly expected to play a central role in AI-driven scientific discovery. Equipped with …
Scoped Verification for Reliable Long-Horizon Agentic Context Evolution under Distribution Shift
arXiv:2607.09175v1 Announce Type: new Abstract: Deployed LLM agents rely on agentic context, the model-external textual control content assembled by an operational harness. In this…
KV-PRM: Efficient Process Reward Modeling via KV-Cache Transfer for Multi-Agent Test-Time Scaling
arXiv:2607.09153v1 Announce Type: new Abstract: Process Reward Models (PRMs) have been proven to be highly effective in guiding test-time scaling (TTS) methods, which significantly…
Neuro-Agentic Control: A Deep Learning-based LLM-Powered Agentic AI Framework for Controlling Security Controls
arXiv:2607.09076v1 Announce Type: new Abstract: Cyberattacks on operational technology are increasingly causing costly downtime and physical damage, exposing the limitations of tra…
ARCANA: A Reflective Multi-Agent Program Synthesis Framework for ARC-AGI-2 Reasoning
arXiv:2607.09059v1 Announce Type: new Abstract: We present ARCANA, a collaborative multi agent framework for solving ARC AGI 2 tasks under strict test time and hardware constraints…
L-MAD: A Systematic Evaluation of Multi-Agent Debate Structures in Legal Reasoning
arXiv:2607.09099v1 Announce Type: new Abstract: While multi-agent debate (MAD) frameworks have shown significant potential in general reasoning, their effectiveness in highly struc…
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading
arXiv:2607.08964v1 Announce Type: new Abstract: AI agents have become capable of autonomously completing short, well-specified tasks. However, existing terminal benchmarks largely …
GATS: Graph-Augmented Tree Search with Layered World Models for Efficient Agent Planning
arXiv:2607.08894v1 Announce Type: new Abstract: Large Language Model (LLM) agents have shown promise in multi-step planning tasks, but existing approaches like LATS (Language Agent…
Same agent tasks, 76% fewer LLM calls – we moved semantic cache inside the graph
Deterministic AI agent orchestration — comment via Discussions & Issues; code changes by maintainers only. - insightitsGit/ChorusGraph…
The authority boundary problem in agent tool calls: who decides what 'no results' means
When a search tool returns nothing, your agent faces a decision: is the query wrong, the index stale, or the backend down? The answer determines what happens next — and most tool-u…