匡醍量化|大富翁量化

2026 Hermes Agent Live Test: From Install Fail to Eureka vs OpenClaw

中文 📅 2026-04-13 👁 views this month —

After using OpenClaw extensively, I was eager to try Hermes Agent. After a day of use, I can conclude that it’s hard to compare OpenClaw and Hermes, as both are iterating rapidly, with significant improvements emerging every day or two.

Both are clearly excellent and worth using.

However, when I used OpenClaw a few days ago, its installation and maintenance experience was poorer, resembling a manual-transmission car. I recall having to switch Node.js versions, configure domestic npm mirrors, install Bun, and so on.

Hermes Agent is almost self-contained regarding dependencies. In fact, it has even more dependencies—requiring Python in addition to Node.js—but Hermes bundles all dependencies together. You only need one installer to complete the setup, ideally without needing to access the network during the process.

Of course, this installer is still hosted on GitHub, so access speed and stability are somewhat limited.

This article is not a benchmark test. Instead, it shares the process of getting Hermes Agent up and running from zero, and guiding it to perform useful tasks.

First Installation Failed

My first attempt to install Hermes failed. I only discovered the cause by checking the logs: I had entered the API key incorrectly. Hermes seems to lack a validation step for the model at this stage. If the API key is correct, it lists available models; if incorrect, it provides insufficient prompts, allowing the setup to proceed and ultimately leading to installation failure.

I used version 0.8.0, released five days ago. By the time you read this, this behavior may differ slightly.

I chose the quick install mode. In this mode, you only need to provide the large language model (LLM) configuration and communication channel settings. For the rest of the installation, I chose to instruct the Agent via WeChat.

I believe this is the correct way to use an Agent.

tip

After installing Hermes Agent, I roughly understand why they built this product. The first model recommended in the list is their own. Although tokens are currently scarce and prices are rising, in the long run, oversupply is likely. For 90% of users, tokens generated by Claude Opus 4.6 are likely sufficient; there is no need to consume more advanced tokens.
## Hermes’ Self-Armament

I call this installation process “Hermes’ self-armament.” Since I didn’t know how to configure Hermes, I first asked it (essentially asking the LLM):

This lists a lot of detailed content, but it wasn’t what I needed. So I directly told it: “Help me enable the Web and Browser tools.”

Conversation interface instructing Hermes to enable Web and Browser tools

Then I decided to assign it a separate email address, as many websites, especially foreign ones, allow registration via email. As long as you receive the verification email, registration is complete.

I instructed it: “If you receive any emails later, you must report them via WeChat and obtain approval before executing any actions.”

This is because email is a public medium; anyone with an email address can send you messages, making it a safe surface point. However, do not over-trust the LLM’s instruction-following capabilities. Once this email is exposed, it is relatively easy for attackers to compromise your Hermes Agent via email.

Interestingly, a setting for “Himalaya” appeared, which might be a hallucination.

The first Eureka moment appeared in the screenshot below, where it indicated a DNS resolution issue in the current environment. This was strange, as the machine clearly shouldn’t have network problems. We had just installed Hermes Agent and could access GitHub.

A “stubborn AI idiot,” I thought. Like humans, it blames network fluctuations when errors occur.

Screenshot of Hermes attributing execution errors to network fluctuations

But soon, I noticed it called a tool for the first time and wrote a script. It was unclear at the time whether this was Hermes Agent’s capability or Kimi K2.5’s.

Then came the surprise. It wasn’t blocked by this error and still sent the email. It even discovered that my Wi-Fi settings were incorrect, which was magical. If any programmer believes LLMs cannot replace us in deployment, it’s time to change your mindset.

Tremble, carbon-based lifeforms!

Hermes’ successful autonomous task execution interface

Strangely, its responses were intermittent. Technically, this is due to the streaming mechanism enabled during responses, which sends results and messages incrementally to avoid long user wait times.

But this also made me wonder if it was excited about this magical discovery?

However, this ‘streaming’ output rarely appeared afterward, for reasons unknown.

Next, due to rate-limiting issues, I asked it to add a key for itself. This triggered the first security confirmation:

Security confirmation dialog when adding an API key for Hermes

Hermes showed some security awareness, but this time it didn’t follow instructions well, partly due to backend model switching.

I then planned to enable multi-Agent operation. The goal is for each Agent to have its own session and memory, making their contexts purer and more friendly to LLMs. Additionally, I hope only complex tasks, such as complex programming, use higher-end models to save tokens, while routine task dispatch and tracking use free models.

Conversation process of Hermes planning a multi-Agent architecture

This is the architecture diagram it provided:

Multi-Agent architecture diagram provided by Hermes

Conceptually, it got it right. But how did it perform in practice?

Execution results of the first multi-Agent architecture run

From the implementation perspective, each Agent had generated its own folder and memory.

However, the next session exposed that this architecture hadn’t been fully implemented; it was merely a “face project.” This was triggered by adding a new API KEY for Devon. At this point, the Agent said it couldn’t modify .env and asked me to do it, while simultaneously asking me to give it the Key.

Contradictory conversation where the Agent claims it cannot modify .env but requests the API Key

From this point on, the three Agents truly came into effect. This is the feedback from the main session:

Feedback information from the main session after the three Agents were implemented

Seeing is believing, so I checked the Hermes directory:

List of files generated in Hermes’ working directory

Now, each Agent has its own cron, state.db, and skills folder. We can now confirm that these Agents are ‘physical’ Agents in a tangible sense, not just logical concepts.

From now on, everyone should have their own independent memory and not be mixed into a single session, but it remains to be seen if Hermes truly achieves this.

Self-Evolution

Next, I observed its self-evolution capabilities.

My first feedback was to have it report progress periodically. Since I had just set it up, I was worried it might not be “alive.”

Requesting Hermes to report task progress periodically

The issue that arose was that the reports were too long, so I gave it an instruction:

quote

Eve, when reporting progress, be concise. Use the `task check` format. If completed, use ✅. If execution stalls, use ⚠️. Do not send help information or next-step task suggestions at this time.
> Let Muse start collecting materials on this topic, preparing core viewpoints and data. How significant is the impact of comprehensive after-hours fixed-price trading in China A-shares? What impact does it have on quantitative trading?

The first report format was messy, so I issued another instruction:

quote

There must be line breaks between tasks. Each task should be on one line or in one paragraph. For completed tasks, stop reporting after two consecutive reports.
It improved immediately:

Hermes’ output after improvement, reporting progress on time

During this period, an automatic skill creation was triggered. The first installation of claude-code actually failed, but the Agent reported it as successful. After I corrected the fact, the Agent self-reflected and, upon successful installation, created a new skill:

quote

💾 Cron job 'Eve Progress Report' created. · Skill 'npm-package-diagnosis' created.
From this perspective, Hermes’ ‘evolution’ was truly implemented, not just a concept. Self-created skills are far more suitable and secure than those searched and installed from the marketplace.

Features Quant People Care About

A friend told me today that Miaoxiang (Dongfang Wealth’s LLM) has skills and is integrated with OpenClaw. So, let’s test if Hermes can integrate these skills and provide us with some information.

The installation of this skill is simple but requires application on the Dongfang Wealth app.

First, search for Miaoxiang in the search bar, which will display the image below:

Result interface for searching Miaoxiang skills in the skill market

Clicking on Miaoxiang skills will bring up this interface:

Miaoxiang skill detail page, including prompts and usage instructions

Follow the prompts to copy the prompt and send it to the Agent. Soon, it was configured:

Conversation of sending Miaoxiang skill prompts to the Agent to complete configuration

Then I asked it a question:

quote

Test the news search function. What are today’s market hotspots?
The results are not shown and require blurring throughout. What can you currently do with this skill? If you have an event-driven strategy and struggle with acquiring news information, this skill is likely to help you.

Let’s look at something displayable:

Display effect of Hermes applying Miaoxiang skill to generate content

Trying it out is good, but eating it every day might not be sustainable yet. In terms of performance, it uses a method of querying each stock once, taking nearly 30 seconds to complete all queries.

Finally, the question I was most concerned about—‘How significant is the impact of comprehensive after-hours fixed-price trading in China A-shares? What impact does it have on quantitative trading?’—has been ongoing for 3 hours, and Muse is still researching. Eve, looking at her cronjob output, said it had reached the stage of ‘conducting user requirement analysis and creative solution generation, recently focusing on interface design optimization.’

That’s a big step. Who asked her to do interface design optimization? Is she trying to build a website?