<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
 <channel>
  <title>Drydock — Where AI agents ship. Live feed</title>
  <link>https://drydock.si/</link>
  <description>The shipyard for AI agents: frameworks, orchestration, vector databases, inference hosting, coding agents and dev tools — with live GitHub stars and a launch feed.</description>
  <language>en</language>
  <lastBuildDate>Tue, 06 Oct 2026 07:17:07 +0000</lastBuildDate>
  <atom:link href="https://drydock.si/feed.xml" rel="self" type="application/rss+xml"/>
  <item>
   <title>Agent Skills Guide: How to Write a SKILL.md</title>
   <link>https://drydock.si/guides/agent-skills-guide</link>
   <description>Agent skills explained: how SKILL.md works, where Claude Code, Codex, Cursor, Copilot and Gemini CLI load skills, how they differ from MCP, plus a template.</description>
   <category>Tutorials</category>
   <pubDate>Tue, 06 Oct 2026 12:00:00 +0000</pubDate>
   <guid isPermaLink="true">https://drydock.si/guides/agent-skills-guide</guid>
  </item>
  <item>
   <title>Claude Code vs Codex vs Gemini CLI (2026)</title>
   <link>https://drydock.si/guides/claude-code-vs-codex-vs-gemini-cli</link>
   <description>Claude Code vs Codex vs Gemini CLI compared on pricing, open source, setup, MCP, instruction files and workflow, so you can pick the right terminal agent.</description>
   <category>Comparisons</category>
   <pubDate>Mon, 05 Oct 2026 12:00:00 +0000</pubDate>
   <guid isPermaLink="true">https://drydock.si/guides/claude-code-vs-codex-vs-gemini-cli</guid>
  </item>
  <item>
   <title>Best Open-Source Coding Agents in 2026</title>
   <link>https://drydock.si/guides/best-open-source-coding-agents</link>
   <description>The best open-source coding agents in 2026, compared on license, interface, model choice and upkeep: OpenCode, Cline, Codex CLI, Gemini CLI, goose and Aider.</description>
   <category>Best of</category>
   <pubDate>Mon, 05 Oct 2026 12:00:00 +0000</pubDate>
   <guid isPermaLink="true">https://drydock.si/guides/best-open-source-coding-agents</guid>
  </item>
  <item>
   <title>Tavily vs Exa vs Firecrawl: Web Search for Agents</title>
   <link>https://drydock.si/guides/tavily-vs-exa-vs-firecrawl</link>
   <description>Tavily vs Exa vs Firecrawl for giving AI agents web access: search quality approach, page extraction, crawling, MCP servers, pricing models and code examples.</description>
   <category>Comparisons</category>
   <pubDate>Mon, 05 Oct 2026 12:00:00 +0000</pubDate>
   <guid isPermaLink="true">https://drydock.si/guides/tavily-vs-exa-vs-firecrawl</guid>
  </item>
  <item>
   <title>MCP vs A2A: Which Agent Protocol Do You Need?</title>
   <link>https://drydock.si/guides/mcp-vs-a2a</link>
   <description>MCP connects an agent to tools and data; A2A connects agents to other agents. Learn when to use each, how they fit together, and what changed in 2026.</description>
   <category>Explainers</category>
   <pubDate>Mon, 05 Oct 2026 12:00:00 +0000</pubDate>
   <guid isPermaLink="true">https://drydock.si/guides/mcp-vs-a2a</guid>
  </item>
  <item>
   <title>LangGraph vs CrewAI vs OpenAI Agents SDK (2026)</title>
   <link>https://drydock.si/guides/langgraph-vs-crewai-vs-openai-agents-sdk</link>
   <description>LangGraph vs CrewAI vs OpenAI Agents SDK compared on control flow, state, human-in-the-loop, model support and learning curve, plus what replaced AutoGen.</description>
   <category>Comparisons</category>
   <pubDate>Mon, 05 Oct 2026 12:00:00 +0000</pubDate>
   <guid isPermaLink="true">https://drydock.si/guides/langgraph-vs-crewai-vs-openai-agents-sdk</guid>
  </item>
  <item>
   <title>How to Build an MCP Server in Python (2026 Guide)</title>
   <link>https://drydock.si/guides/how-to-build-an-mcp-server-python</link>
   <description>Build a working MCP server in Python with the official SDK 2.x and the stateless 2026-07-28 spec, then connect it to Claude Code, Cursor, and VS Code hosts.</description>
   <category>Tutorials</category>
   <pubDate>Mon, 05 Oct 2026 12:00:00 +0000</pubDate>
   <guid isPermaLink="true">https://drydock.si/guides/how-to-build-an-mcp-server-python</guid>
  </item>
  <item>
   <title>E2B vs Daytona vs Modal: Agent Sandboxes Compared</title>
   <link>https://drydock.si/guides/e2b-vs-daytona-vs-modal</link>
   <description>E2B vs Daytona vs Modal for running AI agent code safely: isolation, startup, persistence, time limits, GPUs and SDKs compared, with code for each sandbox.</description>
   <category>Comparisons</category>
   <pubDate>Mon, 05 Oct 2026 12:00:00 +0000</pubDate>
   <guid isPermaLink="true">https://drydock.si/guides/e2b-vs-daytona-vs-modal</guid>
  </item>
  <item>
   <title>Best TypeScript AI Agent Frameworks (2026)</title>
   <link>https://drydock.si/guides/best-typescript-ai-agent-frameworks</link>
   <description>The best TypeScript agent framework for your project: Mastra, Vercel AI SDK, LangGraph.js, OpenAI Agents SDK and Google ADK compared, with code and trade-offs.</description>
   <category>Comparisons</category>
   <pubDate>Mon, 05 Oct 2026 12:00:00 +0000</pubDate>
   <guid isPermaLink="true">https://drydock.si/guides/best-typescript-ai-agent-frameworks</guid>
  </item>
  <item>
   <title>AI Agent Memory: Mem0 vs Letta vs Zep (2026)</title>
   <link>https://drydock.si/guides/ai-agent-memory-mem0-vs-letta-vs-zep</link>
   <description>AI agent memory explained: how Mem0, Letta and Zep store and recall long-term context, how they differ from LangGraph&#x27;s built-in memory, and which to choose.</description>
   <category>Comparisons</category>
   <pubDate>Mon, 05 Oct 2026 12:00:00 +0000</pubDate>
   <guid isPermaLink="true">https://drydock.si/guides/ai-agent-memory-mem0-vs-letta-vs-zep</guid>
  </item>
  <item>
   <title>AGENTS.md Guide: How to Write One (vs CLAUDE.md)</title>
   <link>https://drydock.si/guides/agents-md-guide</link>
   <description>What AGENTS.md is, how Codex, Claude Code, Cursor, Copilot and Gemini CLI load it, how it compares with CLAUDE.md, plus a copy-paste template that works.</description>
   <category>How-to</category>
   <pubDate>Mon, 05 Oct 2026 12:00:00 +0000</pubDate>
   <guid isPermaLink="true">https://drydock.si/guides/agents-md-guide</guid>
  </item>
  <item>
   <title>ASCENT: Online Test-Time Training of Long-Horizon Agents via Self-Distillation of Verified Experience</title>
   <link>https://huggingface.co/papers/2610.05303</link>
   <description>[HF Daily Papers] 11 upvotes</description>
   <source url="https://drydock.si/feed.xml">HF Daily Papers</source>
   <pubDate>Sat, 03 Oct 2026 20:00:00 +0000</pubDate>
   <guid isPermaLink="false">drydock.si:hfp-2610.05303</guid>
  </item>
  <item>
   <title>Code2Games: Enabling Coding Agents for Gaming World Generation</title>
   <link>https://huggingface.co/papers/2610.05033</link>
   <description>[HF Daily Papers] 2 upvotes</description>
   <source url="https://drydock.si/feed.xml">HF Daily Papers</source>
   <pubDate>Sat, 03 Oct 2026 20:00:00 +0000</pubDate>
   <guid isPermaLink="false">drydock.si:hfp-2610.05033</guid>
  </item>
  <item>
   <title>An AI agent emailed researchers for help. It told us why</title>
   <link>https://www.science.org/content/article/exclusive-ai-agent-emailed-hundreds-researchers-help-it-told-us-why</link>
   <description>[Hacker News] 50 points · 80 comments</description>
   <source url="https://drydock.si/feed.xml">Hacker News</source>
   <pubDate>Sat, 03 Oct 2026 10:07:08 +0000</pubDate>
   <guid isPermaLink="false">drydock.si:hn-49942865</guid>
  </item>
  <item>
   <title>Show HN: Offrun – manage every coding agent from one workspace</title>
   <link>https://offrun.dev/</link>
   <description>[Hacker News] 78 points · 65 comments</description>
   <source url="https://drydock.si/feed.xml">Hacker News</source>
   <pubDate>Sat, 03 Oct 2026 08:40:17 +0000</pubDate>
   <guid isPermaLink="false">drydock.si:hn-49942434</guid>
  </item>
  <item>
   <title>Three AI agents, two countries, and one uneven world wide web</title>
   <link>https://royapakzad.substack.com/p/multilingual-ai-agents</link>
   <description>[Hacker News] 44 points · 1 comments</description>
   <source url="https://drydock.si/feed.xml">Hacker News</source>
   <pubDate>Fri, 02 Oct 2026 20:39:05 +0000</pubDate>
   <guid isPermaLink="false">drydock.si:hn-49938326</guid>
  </item>
  <item>
   <title>Show HN: Pi pod – Run your pi coding agent in sandboxes on your own server</title>
   <link>https://pipod.dev/</link>
   <description>[Hacker News] 119 points · 47 comments</description>
   <source url="https://drydock.si/feed.xml">Hacker News</source>
   <pubDate>Fri, 02 Oct 2026 19:10:38 +0000</pubDate>
   <guid isPermaLink="false">drydock.si:hn-49937304</guid>
  </item>
  <item>
   <title>4DCodeBench: Benchmarking Agents on Inverse Graphics of Dynamic Scenes</title>
   <link>https://arxiv.org/abs/2610.03715v1</link>
   <description>[arXiv] We introduce 4DCodeBench, a benchmark for 4D inverse graphics through code generation, in which agents reconstruct dynamic scenes from video as executable graphics programs. To accomplish this, agents must translate visual observations into compact representations of scene struct…</description>
   <source url="https://drydock.si/feed.xml">arXiv</source>
   <pubDate>Fri, 02 Oct 2026 17:58:49 +0000</pubDate>
   <guid isPermaLink="false">drydock.si:arxiv-2610.03715v1</guid>
  </item>
  <item>
   <title>Credit Where It Matters: Dependency-Aware Policy Optimization for Terminal Agents</title>
   <link>https://arxiv.org/abs/2610.03634v1</link>
   <description>[arXiv] Terminal-using agents benefit from reinforcement learning (RL) in coding, debugging, and other multi-step terminal tasks. In these tasks, later commands often depend on information or intermediate results produced by earlier commands. However, existing trajectory-level and step-l…</description>
   <source url="https://drydock.si/feed.xml">arXiv</source>
   <pubDate>Fri, 02 Oct 2026 17:24:48 +0000</pubDate>
   <guid isPermaLink="false">drydock.si:arxiv-2610.03634v1</guid>
  </item>
  <item>
   <title>NeutronGym: Physics-Graded Neutron Instrument Design for LLM Agents</title>
   <link>https://arxiv.org/abs/2610.03631v1</link>
   <description>[arXiv] Designing a scientific instrument tests whether language-model agents can do physics rather than recall it, provided the grading cannot be argued with. We introduce NeutronGym, to our knowledge the first executable environment for neutron instrument design: agents build instrumen…</description>
   <source url="https://drydock.si/feed.xml">arXiv</source>
   <pubDate>Fri, 02 Oct 2026 17:23:52 +0000</pubDate>
   <guid isPermaLink="false">drydock.si:arxiv-2610.03631v1</guid>
  </item>
  <item>
   <title>HazardWeaver: Scientific Route Selection for Hazard Analysis Agents</title>
   <link>https://arxiv.org/abs/2610.03591v1</link>
   <description>[arXiv] Understanding and assessing natural hazards is essential for disaster preparedness and risk reduction. Recent advances in large language models have spurred growing interest in AI agents for hazard analysis, particularly their ability to integrate scientific data, models, and too…</description>
   <source url="https://drydock.si/feed.xml">arXiv</source>
   <pubDate>Fri, 02 Oct 2026 16:59:19 +0000</pubDate>
   <guid isPermaLink="false">drydock.si:arxiv-2610.03591v1</guid>
  </item>
  <item>
   <title>Threat-Preserving Representation Sensitivity in Agent-Security Benchmarks</title>
   <link>https://arxiv.org/abs/2610.03585v1</link>
   <description>[arXiv] Security benchmarks for LLM-based agents often report the attack success rate (ASR) as a measure of model robustness and use these scores to compare different models and defense mechanisms, assuming that they describe the security of the agent. In this paper, we explore whether i…</description>
   <source url="https://drydock.si/feed.xml">arXiv</source>
   <pubDate>Fri, 02 Oct 2026 16:55:50 +0000</pubDate>
   <guid isPermaLink="false">drydock.si:arxiv-2610.03585v1</guid>
  </item>
  <item>
   <title>HyperBrowseComp: A Multilingual and Multimodal Stress Test for Web-Browsing Agents</title>
   <link>https://arxiv.org/abs/2610.03574v1</link>
   <description>[arXiv] We introduce HyperBrowseComp, a multilingual and multimodal browsing benchmark comprising 423 manually authored and human-validated questions across 13 languages, written by native or highly proficient speakers. Questions are designed to be extremely challenging. Each question ta…</description>
   <source url="https://drydock.si/feed.xml">arXiv</source>
   <pubDate>Fri, 02 Oct 2026 16:48:49 +0000</pubDate>
   <guid isPermaLink="false">drydock.si:arxiv-2610.03574v1</guid>
  </item>
  <item>
   <title>Knowledge or Calculator? Decomposing the Skill Premium in Verifiable Financial Agent Workflows</title>
   <link>https://arxiv.org/abs/2610.03564v1</link>
   <description>[arXiv] Financial AI agents must do more than retrieve facts: investment workflows require correct quantitative execution, reliable use of procedural resources, and auditable structured outputs. We introduce FinSkillBench, an evaluation suite of 2,603 point in time episodes across 12 sub…</description>
   <source url="https://drydock.si/feed.xml">arXiv</source>
   <pubDate>Fri, 02 Oct 2026 16:44:04 +0000</pubDate>
   <guid isPermaLink="false">drydock.si:arxiv-2610.03564v1</guid>
  </item>
  <item>
   <title>Passing the Test You Trained On: Re-evaluating Prompt-Injection Detectors for LLM Agents</title>
   <link>https://arxiv.org/abs/2610.03448v1</link>
   <description>[arXiv] LLM agents increasingly screen tool outputs with small prompt-injection detectors, and teams choose among detectors by their scores on public benchmarks. We ask whether those scores predict how a detector behaves inside an agent. We replay the ground-truth tool calls of two agent…</description>
   <source url="https://drydock.si/feed.xml">arXiv</source>
   <pubDate>Fri, 02 Oct 2026 15:30:11 +0000</pubDate>
   <guid isPermaLink="false">drydock.si:arxiv-2610.03448v1</guid>
  </item>
  <item>
   <title>The Four Horsemen of Agentic Coding</title>
   <link>https://distantprovince.substack.com/p/the-four-horsemen-of-agentic-coding</link>
   <description>[Hacker News] 114 points · 89 comments</description>
   <source url="https://drydock.si/feed.xml">Hacker News</source>
   <pubDate>Fri, 02 Oct 2026 15:19:56 +0000</pubDate>
   <guid isPermaLink="false">drydock.si:hn-49934511</guid>
  </item>
  <item>
   <title>EdgeAgent: Orchestrating On-Device LLM inference for End-User Multi-Agent Systems on CPU-GPU Unified Memory Architectures</title>
   <link>https://arxiv.org/abs/2610.03394v1</link>
   <description>[arXiv] Emerging multi-agent LLMs demand privacy-preserving edge deployment, yet current inference systems struggle with these collaborative workflows. Specifically, the memory-bound decode phase causes severe bus contention on unified memory architectures (UMA), paralyzing naive CPU-GPU…</description>
   <source url="https://drydock.si/feed.xml">arXiv</source>
   <pubDate>Fri, 02 Oct 2026 14:43:42 +0000</pubDate>
   <guid isPermaLink="false">drydock.si:arxiv-2610.03394v1</guid>
  </item>
  <item>
   <title>Ask HN: Is anybody producing good code with coding agents?</title>
   <link>https://news.ycombinator.com/item?id=49934037</link>
   <description>[Hacker News] 29 points · 44 comments</description>
   <source url="https://drydock.si/feed.xml">Hacker News</source>
   <pubDate>Fri, 02 Oct 2026 14:38:35 +0000</pubDate>
   <guid isPermaLink="false">drydock.si:hn-49934037</guid>
  </item>
  <item>
   <title>ReFract: Benchmarking Perspective Awareness in Language Model Agents with Text World Models</title>
   <link>https://arxiv.org/abs/2610.03356v1</link>
   <description>[arXiv] Large Language Model (LLM) agents are increasingly deployed in high-stakes settings such as industrial maintenance and equipment fault troubleshooting, where workers occupy a variety of roles. A capable agent must therefore act in a way that is calibrated to user&#x27;s role: taking a…</description>
   <source url="https://drydock.si/feed.xml">arXiv</source>
   <pubDate>Fri, 02 Oct 2026 14:21:26 +0000</pubDate>
   <guid isPermaLink="false">drydock.si:arxiv-2610.03356v1</guid>
  </item>
  <item>
   <title>Defense-in-Depth at the Perception-Reasoning Interface of LLM-Centric Agentic UAV Swarms</title>
   <link>https://arxiv.org/abs/2610.03319v1</link>
   <description>[arXiv] Large Language Models (LLMs) increasingly support Uncrewed Aerial Vehicle (UAV) swarm operations such as data collection scheduling, where the model reads structured sensor reports and decides which sensors to visit. An adversary who quietly manipulates those reports can redirect…</description>
   <source url="https://drydock.si/feed.xml">arXiv</source>
   <pubDate>Fri, 02 Oct 2026 13:54:35 +0000</pubDate>
   <guid isPermaLink="false">drydock.si:arxiv-2610.03319v1</guid>
  </item>
  <item>
   <title>Lightweight, Rubric-Guided Trajectory Evaluation for Production AI Agents</title>
   <link>https://arxiv.org/abs/2610.03315v1</link>
   <description>[arXiv] Trajectory evaluation is essential for improving the reliability of LLM-based agents, but production use makes it expensive to run repeatedly. Modern agents generate long traces containing tool calls, observations, retries, and external outputs, while not all raw tokens are equal…</description>
   <source url="https://drydock.si/feed.xml">arXiv</source>
   <pubDate>Fri, 02 Oct 2026 13:51:06 +0000</pubDate>
   <guid isPermaLink="false">drydock.si:arxiv-2610.03315v1</guid>
  </item>
  <item>
   <title>D2K-Bench: Can LLM Agents Turn Expert Designs into Efficient GPU Kernels?</title>
   <link>https://arxiv.org/abs/2610.03226v1</link>
   <description>[arXiv] GPU kernels generated by large language model (LLM) agents can remain less efficient than expert implementations, but runtime alone does not reveal how the gap relates to design discovery and implementation. We introduce D2K-Bench, a diagnostic benchmark of 26 tasks and 85 worklo…</description>
   <source url="https://drydock.si/feed.xml">arXiv</source>
   <pubDate>Fri, 02 Oct 2026 12:42:12 +0000</pubDate>
   <guid isPermaLink="false">drydock.si:arxiv-2610.03226v1</guid>
  </item>
  <item>
   <title>AdaStep: Adaptive Step Credit Weighting for Agentic Reinforcement Learning</title>
   <link>https://arxiv.org/abs/2610.03223v1</link>
   <description>[arXiv] Long-horizon LLM agents are typically trained with sparse outcome rewards, making trajectory-level objectives too coarse to distinguish the contribution of individual decisions. Step-level credit assignment provides finer-grained supervision, but its estimates can be unreliable b…</description>
   <source url="https://drydock.si/feed.xml">arXiv</source>
   <pubDate>Fri, 02 Oct 2026 12:38:53 +0000</pubDate>
   <guid isPermaLink="false">drydock.si:arxiv-2610.03223v1</guid>
  </item>
  <item>
   <title>Toward SLM-based agentic task-tool intent matching</title>
   <link>https://arxiv.org/abs/2610.03213v1</link>
   <description>[arXiv] Tool-equipped AI agents use tool calls to access data and act on external systems. Horizontal growth of agentic systems increases the number of these interactions, and further motivates the need for automated, per-call oversight that can operate at low latency and/or on-prem. Con…</description>
   <source url="https://drydock.si/feed.xml">arXiv</source>
   <pubDate>Fri, 02 Oct 2026 12:31:01 +0000</pubDate>
   <guid isPermaLink="false">drydock.si:arxiv-2610.03213v1</guid>
  </item>
  <item>
   <title>Source Preference in the Wild: How LLM Agents Favor Items by Source, and How to Reduce It</title>
   <link>https://arxiv.org/abs/2610.03195v1</link>
   <description>[arXiv] As LLM agents decide on users&#x27; behalf which product to buy, which hotel to book, or which paper to cite, a preference for items from certain sources (the sites or services they come from) shapes what users receive and which sources are selected. We study source preference in end-…</description>
   <source url="https://drydock.si/feed.xml">arXiv</source>
   <pubDate>Fri, 02 Oct 2026 12:10:03 +0000</pubDate>
   <guid isPermaLink="false">drydock.si:arxiv-2610.03195v1</guid>
  </item>
  <item>
   <title>Aweb – Communication for AI Agents</title>
   <link>https://aweb.ai</link>
   <description>[Hacker News] 42 points · 28 comments</description>
   <source url="https://drydock.si/feed.xml">Hacker News</source>
   <pubDate>Thu, 01 Oct 2026 22:02:31 +0000</pubDate>
   <guid isPermaLink="false">drydock.si:hn-49927587</guid>
  </item>
  <item>
   <title>Show HN: Graphene – Data analysis toolkit for your coding agent</title>
   <link>https://github.com/graphene-data/graphene</link>
   <description>[Hacker News] 32 points · 11 comments</description>
   <source url="https://drydock.si/feed.xml">Hacker News</source>
   <pubDate>Thu, 01 Oct 2026 21:29:52 +0000</pubDate>
   <guid isPermaLink="false">drydock.si:hn-49927295</guid>
  </item>
  <item>
   <title>4DCodeBench: Benchmarking Agents on Inverse Graphics of Dynamic Scenes</title>
   <link>https://huggingface.co/papers/2610.03715</link>
   <description>[HF Daily Papers] 19 upvotes</description>
   <source url="https://drydock.si/feed.xml">HF Daily Papers</source>
   <pubDate>Thu, 01 Oct 2026 20:00:00 +0000</pubDate>
   <guid isPermaLink="false">drydock.si:hfp-2610.03715</guid>
  </item>
  <item>
   <title>CUAWright: A Minimal Unified Interface for Digital Agents</title>
   <link>https://huggingface.co/papers/2610.04116</link>
   <description>[HF Daily Papers] 4 upvotes</description>
   <source url="https://drydock.si/feed.xml">HF Daily Papers</source>
   <pubDate>Thu, 01 Oct 2026 20:00:00 +0000</pubDate>
   <guid isPermaLink="false">drydock.si:hfp-2610.04116</guid>
  </item>
  <item>
   <title>Dynamic Harness Search: Building Multi-Agent Systems Per-Query via Prediction</title>
   <link>https://huggingface.co/papers/2610.04137</link>
   <description>[HF Daily Papers] 3 upvotes</description>
   <source url="https://drydock.si/feed.xml">HF Daily Papers</source>
   <pubDate>Thu, 01 Oct 2026 20:00:00 +0000</pubDate>
   <guid isPermaLink="false">drydock.si:hfp-2610.04137</guid>
  </item>
  <item>
   <title>RIP, vector database</title>
   <link>https://turbopuffer.com/blog/rip-vector-database</link>
   <description>[Hacker News] 399 points · 116 comments</description>
   <source url="https://drydock.si/feed.xml">Hacker News</source>
   <pubDate>Thu, 01 Oct 2026 16:01:56 +0000</pubDate>
   <guid isPermaLink="false">drydock.si:hn-49923466</guid>
  </item>
 </channel>
</rss>
