Daily AI Roundups
19 June linked agent benchmarks, tool discovery and governance. Artificial Analysis launched AA-Briefcase, a long-horizon benchmark for agentic knowledge work, GLM-5.2 also topped Vals AI's index across legal, finance, proof and coding, and Ollama doubled US GPU capacity for it on Blackwell B300s. Google announced Agentic Resource Discovery, an open spec for agents to find and verify tools, skills and MCP servers, while xAI put Grok on Databricks Agent Bricks and OpenAI Codex added Record and Replay to turn demos into reusable skills. On governance, the White House and Anthropic shifted talks toward frontier AI security rules, and Databricks' playbook catalogued enterprise agent failure cases including stale-policy traces and PII test breaches.
- Agentic AI
- Benchmarks
- Policy & safety
- Dev tooling
- Enterprise AI