OpenClaw's latest update, OpenClaw 2.0, was released after nearly two months of development. This major update, built by 933 contributors, simplifies installation and enhances the browser app, aiming to make OpenClaw more accessible and user-friendly. The release touches every part of OpenClaw, including installation, messaging, memory, skills, models, automations, the browser and native apps, plugins, security, and a long list of fixes.
Reading profile
@codenamev
What @codenamev has been reading — 199 articles, most recently August 31, 2026.
Follow
The paper introduces WikiSkill, a framework that co-evolves agent skills with a persistent knowledge base (wiki). It separates raw execution experience, accumulated knowledge, and executable skills, continuously consolidating experience into the wiki, which subsequent skill updates can build on. Across diverse benchmarks and models, WikiSkill consistently outperforms state-of-the-art skill-evolution methods and improves over no-skill baselines in most model-benchmark settings.
The paper introduces SwarmWorld, a system where initially homogeneous language-model agents self-organize into evolving technological societies without predefined roles. These agents explore environments, process resources, construct artifacts, and develop executable controllers evaluated by a simulator. The study demonstrates that decentralized agents can collaboratively build functional technologies, leading to more resilient technological portfolios compared to isolated search methods. Agents exhibit differentiated behaviors such as exploration, construction, maintenance, and coordination, adapting as the system matures. The research highlights the potential of stigmergic processes in fostering technological evolution within agent societies.
The paper introduces EvoX Genesis, a system that organizes long-horizon software development around a persistent project rather than persistent agents. By representing software as a persistent recursive world, Genesis allows finite-lived agents to propose local changes, with only accepted consequences advancing the persistent version history. The authors demonstrate Genesis's effectiveness by building a Rust-based C compiler from scratch, achieving high performance with low costs, and reimplementing MESA modules with significant speed improvements.
OpenWorker is an open-source, local-first desktop AI assistant designed to automate and complete everyday tasks across various tools and applications. It integrates with services like Slack, email, calendar, and file systems to deliver finished outcomes, such as polished documents, code reviews, and calendar updates, while ensuring user data remains private by operating entirely on the user's device.
The latest edition of Ruby AI News covers significant developments in the Ruby AI ecosystem, including the Rails security team's response to a critical CVE with agent-runnable forensics, advancements in agent technology, and the efficiency of Ruby in agent-based applications.
ArchSpec 1.0 is an architecture linter for Ruby and Rails that allows developers to define components and boundaries in a single file, ensuring that all code changes adhere to the specified architecture. It operates through static analysis, parsing Ruby code to verify compliance with declared rules, without relying on AI. The tool offers presets for various architectural patterns and provides clear error messages when violations occur, facilitating both human and agent-driven development processes.
Daniel Tenner, known as @swombat on X, commented on Anthropic's impressive revenue growth, highlighting the potential for others to achieve similar success.
The grey checkmark on X signifies official government or multilateral organization accounts. Eligibility includes national-level government officials, state-level officials, government organizations, and multilateral organizations. To apply, ensure your profile accurately represents your identity, submit an application form, and confirm your identity through required documents or a government email. Upon approval, your account will be verified.
TimesFM is a pretrained time-series foundation model developed by Google Research for time-series forecasting. It is a decoder-only model designed to handle various time-series data in a zero-shot manner. The latest version, TimesFM 2.5, offers significant improvements over its predecessors, including a reduction in parameters from 500M to 200M, support for up to 16k context length, and the ability to perform continuous quantile forecasting up to a 1k horizon. The model is available on GitHub and can be integrated into various applications, including BigQuery ML, Google Sheets, and Vertex Model Garden.
Raindrop has introduced Signals 2.0, powered by rd-signal-2, a new model pipeline designed to build task-specific binary classifiers from production traces. This advancement aims to achieve high accuracy while significantly reducing costs compared to previous models. The article delves into the challenges of binary classification in AI agents, the development process of rd-signal-2, and its impact on Raindrop's platform. Additionally, it introduces Signal Builder, a platform for training and hosting custom classifiers with Zero Data Retention, catering to environments with strict privacy requirements.
This article discusses the challenges of managing background job queues in multi-tenant Ruby on Rails applications, particularly when a single tenant's large batch of jobs can monopolize resources and delay others. It introduces the 'sidekiq-fair-tenant' gem, which implements fair job prioritization by routing a tenant's jobs to throttled queues after exceeding certain thresholds, ensuring equitable resource distribution among tenants. The article also shares insights from working with Coveralls, a client that monitors test coverage, highlighting the importance of fair job prioritization in multi-tenant systems.
Prime Agent is an open-source reinforcement learning (RL) agent developed by Prime Intellect, designed to enhance coding workflows and manage long-running autonomous tasks. It integrates seamlessly with Prime Intellect's 'verifiers' environments and the 'PRIME-RL' framework, enabling efficient training and evaluation of RL models. The agent's architecture emphasizes modularity and extensibility, allowing for easy customization and scaling. Key features include support for various RL environments, integration with distributed training infrastructures, and a focus on continuous self-improvement through iterative learning processes.
Samuel Williams discusses recent improvements in Ruby's reliability, focusing on handling asynchronous exceptions, managing fiber lifetimes during context switches, and preserving trace events through bytecode optimization. These enhancements aim to ensure that critical information is maintained, allowing for accurate decision-making in scenarios like signal handling, request retries, and server load management.
The Living Library is an interactive platform that allows users to engage in AI-driven conversations with the essence of any book, ancient or modern. By searching for a book and posing a question, users can receive responses that reflect the book's unique perspective. The platform offers a selection of classic texts, such as 'Meditations' by Marcus Aurelius, 'The Art of War' by Sun Tzu, and 'Letters from a Stoic' by Seneca, enabling users to explore these works in a conversational manner.
The Visa Vulnerability Agentic Harness is an open-source project designed to enhance the security of AI agents by providing a robust framework for vulnerability detection and remediation. It integrates seamlessly with existing AI agent infrastructures, offering tools for scanning, analyzing, and fixing security issues within AI-driven applications. This harness aims to automate security assessments, ensuring that AI agents operate securely and efficiently.
FastRuby.io introduces an open-source Claude Code skill designed to automate technical debt audits in Ruby on Rails applications. This tool integrates various existing libraries to streamline the audit process, providing a comprehensive report with a single command. The article outlines the steps to set up and utilize this skill effectively.
Agent Behavior provides a standardized format for defining expected conduct in AI agents across repeated interactions. This framework aids in creating behavior specifications that are both human-readable and machine-processable, ensuring consistency and reliability in agent performance. It emphasizes the importance of clear documentation and version control, integrating seamlessly with existing codebases to maintain alignment between agent behavior and development practices.
Thinking Machines Lab has introduced Inkling, a Mixture-of-Experts transformer model with 975 billion total parameters and 41 billion active parameters. It supports a context window of up to 1 million tokens and has been pretrained on 45 trillion tokens of text, images, audio, and video. Inkling is designed to reason natively over text, images, and audio, balancing cost with performance through efficient and controllable thinking effort. Alongside Inkling, a lighter-weight model, Inkling-Small, with 12 billion active parameters, is also available, achieving strong performance with lower cost and latency. Both models are available for fine-tuning on Tinker, Thinking Machines Lab's training platform.
This article discusses strategies for ensuring AI agents discover, utilize, and ethically interact with your developer tools. It emphasizes the importance of making your tool discoverable to AI agents, facilitating autonomous usage, and implementing safeguards against misuse. The piece also highlights the distinction between baked-in knowledge and live retrieval in AI models, and the necessity of optimizing for both to enhance your tool's visibility.
Helmsman is an adaptive instruction server designed to manage and optimize interactions with various AI coding agents, such as Opus, Sonnet, and Haiku. By dynamically adjusting instructions based on each agent's unique capabilities, costs, and potential failure modes, Helmsman aims to enhance the efficiency and effectiveness of AI-driven code generation processes.
This paper introduces 'agentic reasoning,' a paradigm that redefines large language models (LLMs) as autonomous agents capable of planning, acting, and learning through continuous interaction in dynamic environments. The authors categorize agentic reasoning into three layers: foundational agentic reasoning, self-evolving agentic reasoning, and collective multi-agent reasoning. They also distinguish between in-context reasoning, which scales test-time interaction through structured orchestration, and post-training reasoning, which optimizes behaviors via reinforcement learning and supervised fine-tuning. The paper reviews various agentic reasoning frameworks across real-world applications and benchmarks, including science, robotics, healthcare, autonomous research, and mathematics. It concludes by outlining open challenges and future directions, such as personalization, long-horizon interaction, world modeling, scalable multi-agent training, and governance for real-world deployment.
Claude Code's Auto Dream feature consolidates memory files, pruning stale notes and merging insights, akin to REM sleep for your AI agent.
OneCLI is an open-source credential and policy layer designed to secure AI agents by managing and injecting credentials at the network layer. It operates as a transparent HTTP gateway that intercepts outbound requests, injects credentials from an encrypted vault, and enforces access policies, ensuring agents never possess raw API keys. This approach mitigates risks associated with credential exposure and unauthorized access.
Ctrl+Shft offers a comprehensive solution for managing dotfiles tailored for AI coding agents, addressing common challenges such as context degradation, instruction drift, secret exposure, and rule pollution. By providing a unified repository, it ensures consistent and secure environments across different machines, enhancing the efficiency and reliability of AI-driven development workflows.
RubyCritic is a tool designed to analyze and report on the quality of Ruby codebases. By integrating various static analysis tools, it provides a comprehensive overview of code health, highlighting areas that may require attention. This approach aids developers in maintaining high-quality, maintainable code.
Next Rails is a toolkit designed to assist developers in upgrading their Ruby on Rails applications. It offers a set of tools and guidelines to streamline the upgrade process, ensuring compatibility with the latest Rails versions and best practices. By leveraging Next Rails, developers can enhance the performance, security, and maintainability of their applications, facilitating smoother transitions to newer Rails releases.
ruby-next is a transpiler and collection of polyfills designed to support the latest and upcoming Ruby features in older versions and alternative implementations. It enables developers to utilize modern Ruby syntax and APIs, such as pattern matching and Kernel#then, in environments like Ruby 2.5 or mruby. This tool is particularly beneficial for gem maintainers aiming to write code compatible with both current and legacy Ruby versions, as well as for developers eager to experiment with new features without waiting for official releases.
This tutorial provides a comprehensive guide on integrating Hermes Agent with OpenRouter, covering setup procedures, model selection, and routing configurations. It emphasizes the importance of selecting models with at least 64K context tokens to ensure optimal performance and discusses various routing modes like `openrouter/auto` and `openrouter/pareto-code` for specific use cases. The guide also details the configuration of fallback chains and auxiliary-model offloading within the `~/.hermes/config.yaml` file, offering practical insights for users to effectively deploy and manage Hermes Agent with OpenRouter.
Lilian Weng's article explores the concept of harness engineering in AI, emphasizing its role in facilitating recursive self-improvement (RSI). She discusses how harnesses—systems surrounding AI models—are crucial for orchestrating execution, managing context, and enabling models to improve autonomously. The article delves into design patterns for harnesses, optimization strategies, and future challenges in the field. Weng also provides an appendix with useful benchmarks for evaluating AI agents.
Shepherd is a Python-based framework that transforms an agent's execution into a reversible, Git-like trace, enabling meta-agents to observe, fork, replay, and revert any run. This approach facilitates efficient supervision, optimization, and training of agents. The system records every agent-environment interaction as a typed event, allowing for precise control over agent behavior. Applications include runtime intervention, counterfactual optimization, and tree-search reinforcement learning, demonstrating significant improvements in performance and efficiency.
This paper introduces 'agentic reasoning,' a paradigm that redefines large language models (LLMs) as autonomous agents capable of planning, acting, and learning through continuous interaction in dynamic environments. The authors categorize agentic reasoning into three layers: foundational agentic reasoning, self-evolving agentic reasoning, and collective multi-agent reasoning. They also distinguish between in-context reasoning, which scales test-time interaction through structured orchestration, and post-training reasoning, which optimizes behaviors via reinforcement learning and supervised fine-tuning. The paper reviews various agentic reasoning frameworks across real-world applications and benchmarks, including science, robotics, healthcare, autonomous research, and mathematics. It concludes by outlining open challenges and future directions, such as personalization, long-horizon interaction, world modeling, scalable multi-agent training, and governance for real-world deployment.
ZeroLang is a programming language developed by Vercel Labs, designed specifically for building AI agents. It offers a streamlined syntax and robust features tailored for agent development, enabling developers to create intelligent, autonomous systems efficiently. The language emphasizes simplicity and performance, making it an ideal choice for AI applications that require quick development cycles and reliable execution.
Justin Poehnelt discusses the necessity of redesigning command-line interfaces (CLIs) to prioritize AI agents as primary users, emphasizing the importance of machine-readable outputs, schema introspection, and safety measures to enhance agent interaction and efficiency.
Zero is a programming language designed with agents as primary users from the outset. It emphasizes ease of learning, deterministic inspection and repair, a comprehensive standard library, and explicitness to ensure clear and straightforward task execution. The language aims to be learnable on demand, with a small surface area and regular syntax, and to provide deterministic repair loops through structured diagnostics and repair plans.
Dramabox is an expressive text-to-speech (TTS) model developed by Resemble AI, built upon Lightricks' LTX-2.3 audio branch. It enables users to generate speech with controlled speaker identity, emotion, delivery style, and paralinguistic features like laughs and pauses. By providing a 10-second voice reference, users can clone a target voice's timbre. The model is available on Hugging Face under the LTX-2 Community License.
AVB (@neural_avb) has recently shared several insights on AI advancements and software development practices. Notably, AVB highlighted a comprehensive article detailing the evolution of Convolutional Neural Networks (CNNs) and their performance in the ImageNet competitions of the mid-2010s, covering architectures like LeNet, AlexNet, VGG, Inception, ResNet, DenseNets, and SENet. Additionally, AVB discussed the challenges faced by OpenAI's ChatGPT, particularly its limitations in handling code execution and the impact of these constraints on user experience. Furthermore, AVB emphasized the importance of specification-driven development, advocating for systems that are externally controlled through JSON/YAML abstractions and internally structured with clear module specifications and type validation.
This study introduces a multimodal latent diffusion model that synthesizes volumetric magnetic resonance imaging (MRI) and tabular clinical data within a shared latent space using cross-attention mechanisms. This approach enables coherent joint representation learning, facilitating improved integration and analysis of multimodal medical data.
OpenAI's team developed a software product entirely without manually written code, utilizing Codex to generate all aspects, including application logic, tests, and documentation. This approach significantly accelerated development, completing the project in about one-tenth the time compared to traditional methods. The experiment highlighted the evolving role of engineers, focusing on designing environments and feedback loops to enable Codex agents to perform reliably. Key lessons included redefining engineering roles, enhancing application readability, and understanding the implications of agent-generated code. The team continues to explore how to maximize human time and attention in this new paradigm.
Flue is an open-source framework designed to facilitate the development and deployment of sandboxed agents. It provides a structured environment for creating, testing, and managing agents in isolated settings, ensuring security and stability during development. The framework is built with modularity in mind, allowing developers to customize and extend its components to suit various use cases. Flue is actively maintained and encourages contributions from the community to enhance its capabilities and support a wide range of applications.
Amazon's Strands Agents introduces Agent SOPs, a standardized markdown format for defining AI agent workflows in natural language. This approach balances control and flexibility, enabling teams to create reusable, shareable workflows that guide agent behavior consistently across different AI systems and teams. By combining structured guidance with the adaptability of AI agents, Agent SOPs address challenges like inconsistent behavior and complex prompt engineering, facilitating more reliable and efficient AI agent development.
The article introduces the 'Intent Layer,' a context engineering system designed to enhance AI agents' performance on large codebases by embedding a team's institutional knowledge directly into the codebase. It discusses the challenges agents face due to limited context and how the Intent Layer addresses these by providing hierarchical, token-efficient context through 'Intent Nodes.' The piece also outlines the process of building and maintaining the Intent Layer, emphasizing its benefits in improving agent efficiency and reducing maintenance overhead.
The AI Coding platform at Leaflet Pub provides an interactive web-based environment designed to assist developers with coding tasks through AI-powered tools. It offers functionalities to generate code, debug, optimize, and provide explanations for various programming problems, aiming to enhance productivity and learning for users. The interface includes features like code input areas, output display, and model interaction, enabling seamless AI-driven coding support. This platform leverages advanced AI models to facilitate developers in writing, understanding, and refining code efficiently.
The X Developer Platform Status page provides real-time updates on the operational status of X's developer services, including the X API v2, GNIP Enterprise API, and Developer Console. It also lists recent incidents and their resolutions.
In this essay, Paul Graham presents insights from Richard Hamming's lecture on conducting impactful research. Hamming emphasizes the importance of curiosity, courage, and the willingness to tackle significant problems. He discusses the necessity of periodically shifting focus to prevent stagnation and the value of making one's work accessible for others to build upon. Hamming also highlights the role of self-management in overcoming personal faults to achieve great work.
PostHog shares lessons learned from two years of developing AI agents, emphasizing the importance of considering whether to build a custom AI agent or provide access to existing agents through an MCP server. They discuss the challenges of creating a unique agent harness and the significance of leveraging existing solutions. The article also highlights the value of integrating product context into AI agents to enhance their effectiveness and the necessity of establishing observability and evaluation mechanisms from the outset to monitor and improve AI agent performance.
This article provides a practical guide on managing sessions, context, and compaction in Claude Code, especially with the new 1 million token context window. It discusses the impact of session management on results and offers strategies for effective usage.
Anthropic has launched Claude Design, a new product that enables users to collaborate with Claude to create polished visual work such as designs, prototypes, slides, and more. Powered by Claude Opus 4.7, Claude Design is available in research preview for Claude Pro, Max, Team, and Enterprise subscribers.
The article discusses how the integration of AI into software development impacts Agile methodologies. It emphasizes that while the core principles of Agile—such as communication loops and short feedback cycles—remain unchanged, the roles within these processes are evolving. AI agents are increasingly taking on the role of authors, with human developers acting more as editors or directors. This shift necessitates adjustments in Agile practices, particularly in managing the volume and complexity of changes introduced by AI, and underscores the importance of human-driven reviews to maintain shared understanding within development teams.
An analysis of 26 advanced language models reveals a convergence towards a 'contemplative essayist' style, with distinct postures maintained by each lab. Notably, Anthropic's models exhibit introspective hedging, while Google's Gemini models employ mechanistic language. The study also highlights shared lexical patterns across different labs, suggesting potential information leakage. These insights underscore the evolving nature of AI-generated content and the influence of training methodologies.
This page is also readable by software — the same list is available as RSS, Markdown, and llms.txt for feed readers and AI agents.