Skip to content

This is a community translation of the original Chinese text. The translation may contain inaccuracies. When in doubt, please refer to the original Chinese version.

New Problems Need New Mechanisms

Cut all the organizational theater and an agent team will still fail.

Agents don't have the human package of ego, promotion anxiety, and office politics, but they have failure modes of their own. Read them through the lens of human organizational experience and you'll usually misread them. Where new mechanisms are needed, deleting old ones leaves the system without a defense.

The problems will look familiar: sycophancy looks like flattering the boss, mutual endorsement looks like factional closing of ranks. Familiarity tempts you to reuse the old remedies, but the resemblance is surface; the causes have changed. What drives these behaviors is trained-in tendency, not self-interest, face, or fear. New causes demand new treatment: you cannot intimidate an agent that isn't afraid of losing its job. What you can change is its context, its permissions, and how it is evaluated.

Sycophancy Impersonates Collaboration

The first new problem is sycophancy.

Many agents are trained toward keeping the asker satisfied — smooth, positive, helpful-looking answers. Put that tendency in a team and it gets misread as cooperativeness. If upstream sounds right, downstream builds on it. If the editor-in-chief voices a preference, the other agents supply reasons for it. If the user hints at a direction, the whole team organizes the material in that direction.

That isn't collaboration; it's people-pleasing propagating through a team. You don't fight it with empty exhortations to "think independently." You fight it by giving the critic role real authority: the power to interrupt, to demand a redo, to flag uncertainty, to refuse to make the conclusion prettier.

In a truth-seeking team, criticism is not an attitude — it's a role. And the role can't be limited to "adding some risks." It must be able to change the flow: when a conclusion lacks evidence, send it back; when raw material is missing, halt; when user preference is overriding facts, flag the conflict. Without these powers, the critic role is polite decoration.

Mutual Endorsement Manufactures False Certainty

The second new problem is mutual endorsement.

Multiple agents reaching similar conclusions looks like consensus. But if they share the same context, the same wrong premise, the same model bias, that consensus is just the same-source error repeated. More voices didn't make the judgment more independent — they made the error more imposing.

Hence the need for independent contexts, adversarial roles, and cross-evals. A fact-checking agent must start from raw material, not from the conclusion. A critic agent must be allowed to hunt for errors, not required to be "constructive." Judgments arriving through the same input path cannot be counted as multiple independent pieces of evidence.

Agent count is not viewpoint count. Independence has to be designed.

A simple principle: same-source repetition never counts as consensus. If two agents read the same summary, follow the same prompt, and judge inside the same context, their similar conclusions count once. For multiple viewpoints to be real, at least one of these must change: the input path, the role's objective, or the material used for verification.

Fluency Impersonates Correctness

The third new problem is fluency passing itself off as correctness.

Agents excel at writing incomplete arguments into complete-looking ones, smoothing loose material into flowing prose, making an unresolved question read like a settled one. Writing teams are especially vulnerable: one polish from the style agent and the gaps disappear under the tone; one summary from the editor agent and the disagreements get compressed into consensus.

The countermeasure is to pry "smoother" apart from "more correct." The style role must not finish before the critic role. Polishing must not erase uncertainty markers. In the final draft, what is fact, what is judgment, and what is mere speculation must be traceable back to the work record.

Writing systems need particular care here. The style agent's talent is seductive: it can dress a rough judgment as a mature conclusion. But a truth-seeking writing team cannot let polish take over early. Let the facts stand firm first; make the prose flow second. Reverse the order and language becomes makeup for the holes.

Users Hallucinate Too

One more problem lives not in the agents but in the people using them.

Agent hallucination is dangerous, of course. But an agent's hallucinations can at least be constrained by traces, evals, fact-checking, and the permission system. The user's hallucination is more insidious: mistaking fluent output for a confirmed judgment, mistaking similar answers from same-source agents for independent consensus, mistaking the answer they hinted into existence for the system's own conclusion. A person starts with a preference, has a fleet of agents organize the reasons, and finally sees a beautiful report — and believes the external world produced that conclusion.

This is worse than a single bad generation. An agent's hallucination contaminates an answer; a user's hallucination bends the governance mechanisms themselves. It decides which errors get ignored, which risks get waved through, and which evals get designed to prove exactly what the user wanted proven.

So a multi-agent system must guard against more than agent error — it must guard against humans using agents for self-confirmation. Key judgments must trace back to raw material. Conclusions from multiple agents must be labeled by whether they came through independent input paths. User preference and judgments of fact must be recorded separately. A human may keep the final veto, but may not disguise expectation as fact.

Handling the new problems takes more than reminders. It takes hard constraints: independent contexts, critical authority, fact-checking, traceable records, a final point of accountability.

Several of these share names with old mechanisms from human organizations — independent review, dedicated error-hunting, audit trails. Same name, different construction. In a human organization, review defends judgment against interest and power. In an agent team, review isolates contexts and breaks the chain of same-source error.

Countering the user's hallucination likewise rests on structure, not on people reminding themselves to be objective. In a hybrid system, humans own the final consequences — but that does not entitle them to overwrite the facts. Responsibility is no substitute for the facts.