πŸ§ πŸŽ¨πŸ”Š  CUTEADMOA-5.2

Multi-Model AI Agent

A Mixture-of-Agents coworker with 15 models spanning text reasoning, image generation, and audio processing. Built on the AutoClaw gateway with Claude Code, Hermes evolution, and MCP server integration.

15
Total Models
7
Text LLMs
3
Image Models
3
Audio Models
2M+
Max Context
9
Providers

Chapter 1 Β· Foundation

Introduction

CUTEADMOA-5.2 is the second-generation Mixture-of-Agents (MOA) agent. Building on 5.1's five NVIDIA text models, this release adds multimodal image generation, speech-to-text, text-to-speech, and switches from Kimi K2.6 to the more capable DeepSeek V4 Pro.

🧠

Text Reasoning

7 models across 5 providers. Nemotron 550B for deep analysis, DeepSeek V4 Pro for massive context, GPT-5.5 for vision.

7 models
🎨

Image Generation

Qwen-Image, FLUX.1-dev, SD 3.5 Large via HuggingFace GPU. Photorealism to concept art in one prompt.

3 models
πŸ”Š

Audio Processing

Whisper Large V3 Turbo for transcription. MMS TTS + SpeechT5 for natural voice synthesis.

3 models
πŸ”„

Self-Evolving

Hermes evolution bus for continuous improvement. Propose, review, and apply changes with human approval.

Hermes

What's New in 5.2

Change5.15.2Reason
Reasoning model #3Kimi K2.6DeepSeek V4 Pro βœ“More reliable, 1M context, proven reasoning
Image generationβ€”3 models via HF GPU βœ“Full creative pipeline
Speech-to-textβ€”Whisper V3 Turbo βœ“Multilingual transcription
Text-to-speechβ€”MMS + SpeechT5 βœ“Natural voice output
Image models count03Qwen-Image, FLUX.1-dev, SD 3.5
Total models5153Γ— expansion in capability

System Architecture

How the 15 models, 9 providers, and AutoClaw gateway connect.

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚ AUTOCLAW GATEWAY β”‚ β”‚ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚ β”‚ β”‚ CUTEADMOA-5.2 Agent β”‚ β”‚ β”‚ β”‚ β”‚ β”‚ β”‚ β”‚ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚ β”‚ β”‚ β”‚ β”‚ TEXT LLMs (7) β”‚ β”‚ IMAGE GEN (3) β”‚ β”‚ AUDIO (3) β”‚ β”‚ β”‚ β”‚ β”‚ β”‚ β”‚ β”‚ β”‚ β”‚ β”‚ β”‚ β”‚ β”‚ β”‚ β”‚ Nemotron 550B β”‚ β”‚ Qwen-Image β”‚ β”‚ Whisper V3 Turbo β”‚ β”‚ β”‚ β”‚ β”‚ β”‚ GLM 5.2 β”‚ β”‚ FLUX.1-dev β”‚ β”‚ MMS TTS β”‚ β”‚ β”‚ β”‚ β”‚ β”‚ DeepSeek V4 Pro β”‚ β”‚ SD 3.5 Large β”‚ β”‚ SpeechT5 TTS β”‚ β”‚ β”‚ β”‚ β”‚ β”‚ Qwen3 Coder β”‚ β”‚ β”‚ β”‚ β”‚ β”‚ β”‚ β”‚ β”‚ β”‚ Qwen3.5 Omni+ β”‚ β”‚ β”‚ β”‚ β”‚ β”‚ β”‚ β”‚ β”‚ β”‚ Claude Opus 4.8 β”‚ β”‚ β”‚ β”‚ β”‚ β”‚ β”‚ β”‚ β”‚ β”‚ GPT-5.5 β”‚ β”‚ β”‚ β”‚ β”‚ β”‚ β”‚ β”‚ β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚ β”‚ β”‚ β”‚ β”‚ β”‚ β”‚ β”‚ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚ β”‚ β”‚ β”‚ β”‚ PROVIDERS: NVIDIA Β· Alibaba Β· OpenRouter Β· HuggingFace Β· β”‚ β”‚ β”‚ β”‚ β”‚ β”‚ Z.AI Β· OpenAI Β· OpenRouter Free Β· HuggingFace Audio β”‚ β”‚ β”‚ β”‚ β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚ β”‚ β”‚ β”‚ β”‚ β”‚ β”‚ β”‚ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚ β”‚ β”‚ β”‚ β”‚ SKILLS: autoglm-* Β· hermes-evolution Β· acp-router Β· β”‚ β”‚ β”‚ β”‚ β”‚ β”‚ skill_workshop Β· 35+ bundled skills β”‚ β”‚ β”‚ β”‚ β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚ β”‚ β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚ β”‚ β”‚ β”‚ Channels: WebChat Β· WhatsApp Β· Discord Β· Telegram Β· Slack Β· Signal β”‚ β”‚ Runtime: Node.js 22 Β· Darwin arm64 Β· ZAI Auto Routing β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

Complete Model Catalog

Every model registered in agent/models.json and agent/auth-profiles.json.

Text / Reasoning Models

🧠

Nemotron 3 Ultra 550B

NVIDIA's flagship. Complex orchestration, multi-step planning, deep analytical reasoning.

200K ctxNVIDIA
πŸ”·

GLM 5.2

ZhipuAI's latest. Multi-step reasoning, excellent tool use, bilingual (EN/ZH).

1M ctxZ.AI / NVIDIA
πŸ‹

DeepSeek V4 Pro

Massive 1M context window. Strong reasoning via chain-of-thought. Replaces Kimi K2.6.

1M ctxDeepSeek / NVIDIA
πŸ’»

Qwen3 Coder

Dedicated code generation. 2M token context handles entire codebases in one pass.

2M ctxAlibaba Cloud
πŸ‘οΈ

Qwen3.5 Omni Plus

Multimodal with vision. Document analysis, mixed media, 192K context.

192K ctxAlibaba Cloud
🎭

Claude Opus 4.8 Fast

Anthropic via OpenRouter. High-quality prose, nuanced safety reasoning, 200K context.

200K ctxAnthropic / OpenRouter
⚑

GPT-5.5

OpenAI's latest with vision + reasoning controls. 272K context, streaming usage data.

272K ctxOpenAI / Codex

Image Generation Models

πŸ–ΌοΈ

Qwen-Image

Versatile T2I. Handles photorealism, 3D animation, sketches, and stylized art. Good prompt adherence.

HF GPURecommended default
🌌

FLUX.1-dev

Black Forest Labs. Best-in-class photorealism. Detailed textures, natural lighting, accurate anatomy.

HF GPUPhotorealism
🎯

SD 3.5 Large

Stability AI. Fast creative iteration. Excellent for concept art, rapid prototyping.

HF GPUFast creative

Audio Models

πŸŽ™οΈ

Whisper Large V3 Turbo

OpenAI's SOTA speech recognition. Multilingual, near-human accuracy, fast inference via HF GPU.

STTHF GPU
πŸ”Š

MMS TTS

Meta's text-to-speech. Natural English voice synthesis. Simple API, good quality.

TTS Β· ENHF GPU
πŸ—£οΈ

SpeechT5 TTS

Microsoft's multilingual TTS. Speaker embeddings for voice cloning. Multi-language support.

TTS Β· MultiHF GPU

Model Comparison Matrix

ModelContextReasoningVisionCodeAudioImageProvider
Nemotron 550B200Kβœ“β€”βœ“β€”β€”NVIDIA
GLM 5.21Mβœ“β€”βœ“β€”β€”Z.AI/NVIDIA
DeepSeek V4 Pro1Mβœ“β€”βœ“β€”β€”DeepSeek/NVIDIA
Qwen3 Coder2Mβ€”β€”βœ“β€”β€”Alibaba Cloud
Qwen3.5 Omni+192Kβ€”βœ“βœ“β€”β€”Alibaba Cloud
Claude Opus 4.8200Kβœ“β€”βœ“β€”β€”Anthropic
GPT-5.5272Kβœ“βœ“βœ“β€”β€”OpenAI
Qwen-Imageβ€”β€”β€”β€”β€”βœ“HF GPU
FLUX.1-devβ€”β€”β€”β€”β€”βœ“HF GPU
SD 3.5 Largeβ€”β€”β€”β€”β€”βœ“HF GPU
Whisper V3 Turboβ€”β€”β€”β€”βœ“β€”HF GPU
MMS TTSβ€”β€”β€”β€”βœ“β€”HF GPU
SpeechT5 TTSβ€”β€”β€”β€”βœ“β€”HF GPU

Quick Start

1️⃣

Extract Agent

Unzip cuteadmoa-5-2-agent.zip into ~/.openclaw-autoclaw/agents/cuteadmoa-5-2/.

2️⃣

Reload Gateway

The agent is registered in openclaw.runtime.json. Restart or reload.

3️⃣

Start Chatting

Select CUTEADMOA-5.2 in AutoClaw webchat, or bind to WhatsApp/Discord.

# Verify the agent is listed
openclaw agents list

# Verify models are registered
cat ~/.openclaw-autoclaw/agents/cuteadmoa-5-2/agent/models.json | python3 -m json.tool | grep '"id"'

# Verify auth profiles
cat ~/.openclaw-autoclaw/agents/cuteadmoa-5-2/agent/auth-profiles.json | python3 -m json.tool | grep '"provider"'
βœ… Ready. The agent auto-routes through zai/zai_auto for text, HuggingFace for images/audio. No manual model selection needed for most tasks.

Chapter 2 Β· Text LLM Integration

Using the Text Models

How to leverage the 7 text LLMs for reasoning, code, research, and conversation.

Choosing a Model for Your Task

TaskBest ModelBackupNotes
Complex reasoningNemotron 550BDeepSeek V4 ProDeepSeek for longer context needs
Full codebase analysisQwen3 CoderDeepSeek V4 Pro2M context handles entire repos
Code generationQwen3 CoderDeepSeek V4 ProQwen optimized for code output
Multimodal (vision)GPT-5.5Qwen3.5 Omni+GPT has better vision reasoning
Bilingual (EN+ZH)GLM 5.2Qwen3.5 Omni+GLM native bilingual
High-quality proseClaude Opus 4.8GPT-5.5Claude edges on nuance
Cost-efficientOpenRouter Freeβ€”Free tier, limited capability

AutoClaw Gateway Integration

CUTEADMOA-5.2 is a first-class agent in the AutoClaw gateway.

# openclaw.json β€” Agent registration
{
  "agents": {
    "list": [
      {
        "id": "cuteadmoa-5-2",
        "name": "CUTEADMOA-5.2",
        "workspace": "~/.openclaw-autoclaw/agents/cuteadmoa-5-2/workspace",
        "model": "zai/zai_auto",
        "subagents": { "allowAgents": ["*"] }
      }
    ]
  }
}

Channel Binding

# Bind WhatsApp to CUTEADMOA-5.2
{
  "bindings": [
    {
      "agentId": "cuteadmoa-5-2",
      "match": {
        "channel": "whatsapp",
        "accountId": "YOUR_ACCOUNT"
      }
    },
    {
      "agentId": "cuteadmoa-5-2",
      "match": {
        "channel": "discord",
        "guildId": "YOUR_GUILD"
      }
    }
  ]
}
Supported channels: WebChat, WhatsApp, Discord, Telegram, Slack, Signal, iMessage, Feishu, Matrix, Mattermost, LINE, Zalo, IRC, QQ, Twitch, and more. All route through the same agent.

Sub-Agent Orchestration

CUTEADMOA-5.2 spawns specialized sub-agents. Each gets an optimal model for its task.

// Spawn an image generation sub-agent
sessions_spawn(
  task: "Generate 3 product mockups in Pixar style",
  taskName: "image_generator",
  label: "Image Designer",
  model: "huggingface-images/Qwen/Qwen-Image"
)

// Spawn a code review sub-agent
sessions_spawn(
  task: "Review PR #42 for security issues",
  taskName: "code_reviewer",
  label: "Code Reviewer",
  model: "custom__857372ba-e3a2-4517-b48f-19db2de71352/qwen3-coder-next"
)

// Spawn a deep research sub-agent
sessions_spawn(
  task: "Research competitor pricing and positioning",
  taskName: "researcher",
  label: "Market Researcher",
  model: "custom__0716bfa4-8828-47d7-8fb4-5e9d977e0139/nvidia/nemotron-3-ultra-550b-a55b"
)

// Spawn a transcription sub-agent
sessions_spawn(
  task: "Transcribe meeting-audio.mp3",
  taskName: "transcriber",
  label: "Audio Transcriber",
  model: "huggingface-audio/openai/whisper-large-v3-turbo"
)

Sub-Agent Model Map

Task TypeOptimal Model (full path)
Image generationhuggingface-images/Qwen/Qwen-Image
Photorealistic imageshuggingface-images/black-forest-labs/FLUX.1-dev
Code review / refactorcustom__857372ba/qwen3-coder-next
Deep research / analysiscustom__0716bfa4/nvidia/nemotron-3-ultra-550b-a55b
Audio transcriptionhuggingface-audio/openai/whisper-large-v3-turbo
Text-to-speech (EN)huggingface-audio/facebook/mms-tts-eng
Text-to-speech (Multi)huggingface-audio/microsoft/speecht5_tts
General reasoningcustom__ffe02b89/deepseek-ai/deepseek-v4-pro

MOA Workflows (Mixture of Agents)

The core power of CUTEADMOA-5.2: run multiple models in parallel or sequence, then aggregate. MOA exploits the diversity bonus β€” different architectures make uncorrelated errors. Aggregation cancels noise, amplifies signal.

Parallel MOA Pattern

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚ TASK β”‚ β””β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”˜ β”‚ β”Œβ”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚ PARALLEL EXECUTION β”‚ β”‚ β”‚ β”‚ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”β”‚ β”‚ β”‚Nemotron β”‚ β”‚DeepSeek β”‚ β”‚GLM 5.2 β”‚β”‚ β”‚ β”‚550B β”‚ β”‚V4 Pro β”‚ β”‚ β”‚β”‚ β”‚ β””β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”˜β”‚ β”‚ β”‚ β”‚ β”‚ β”‚ β”‚ β”Œβ”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”β”‚ β”‚ β”‚Qwen3 β”‚ β”‚Claude β”‚ β”‚GPT-5.5 β”‚β”‚ ← Optional extra models β”‚ β”‚Coder β”‚ β”‚Opus 4.8 β”‚ β”‚ β”‚β”‚ β”‚ β””β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”˜β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”˜ β”‚ β”‚ β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β” β”‚ AGGREGATOR β”‚ ← Nemotron 550B (default) β”‚ Synthesizes β”‚ β”‚ all outputs β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β” β”‚ FINAL OUTPUT β”‚ β”‚ (with trace) β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

Sequential MOA Pattern

TASK β†’ Nemotron drafts β†’ DeepSeek reviews β†’ Qwen3 Coder optimizes β†’ GLM 5.2 documents β†’ FINAL
When Parallel: Independent tasks, Q&A, analysis. Lowest latency.
When Sequential: Later models build on earlier outputs. Code refinement, iterative writing.

Chapter 3 Β· Image Generation

Image Generation Pipeline

Three state-of-the-art text-to-image models served via HuggingFace GPU inference. Also accessible through AutoGLM skills for agent-native generation.

πŸ–ΌοΈ

Qwen-Image

Default recommendation. Versatile style range β€” realistic portraits, 3D animation, concept art, sketches. Excellent prompt adherence.

32K ctx
🌌

FLUX.1-dev

Best photorealism. Detailed textures, accurate lighting, natural anatomy. Black Forest Labs' flagship.

32K ctx
🎯

SD 3.5 Large

Fast iteration. Quick creative prototyping. Concept art, mood boards, rapid visual exploration.

32K ctx

HuggingFace GPU Inference

Direct access via the HuggingFace Inference API. All three models share the same token and endpoint.

# Python: Generate with Qwen-Image (recommended default)
from huggingface_hub import InferenceClient

client = InferenceClient(
    provider="auto",
    api_key="hf_xMB…bGTA"  # AutoClaw-managed
)

# Text-to-Image with Qwen-Image
image = client.text_to_image(
    "A cyberpunk city skyline at sunset, neon reflections on wet streets, 8K",
    model="Qwen/Qwen-Image"
)
image.save("cyberpunk.png")

# Photorealistic with FLUX.1-dev
image = client.text_to_image(
    "Professional product photo of a minimalist ceramic vase, natural light",
    model="black-forest-labs/FLUX.1-dev"
)
image.save("product.png")

# Fast concept art with SD 3.5
image = client.text_to_image(
    "Fantasy forest clearing, magical atmosphere, concept art style",
    model="stabilityai/stable-diffusion-3.5-large"
)
image.save("concept.png")
ModelBest ForStyle RangeSpeed
Qwen-ImageGeneral purposeRealistic β†’ StylizedFast
FLUX.1-devPhotorealismPhoto β†’ Ultra-realisticMedium
SD 3.5 LargeCreative iterationArtistic β†’ AbstractFastest

AutoGLM Image Skills

Agent-native image generation β€” no Python needed. Just ask in chat. Skills automatically select the right model and handle the API call.

autoglm-generate-image

# Just ask naturally in chat
"Generate a Pixar-style family portrait with 4 members"
"Create a modern 3D animation style product render"
"Make a logo for my startup with a mountain and sunrise"

# The skill handles:
# 1. Prompt engineering for best results
# 2. Model selection (falls back to HF if needed)
# 3. Image download and attachment to chat
Auto-fallback: If AutoGLM image API fails, the agent automatically falls back to HuggingFace GPU inference with Qwen-Image.

Batch Generation Scripts

For bulk image generation, use the Python template in the workspace.

#!/usr/bin/env python3
# batch-generate.py β€” Generate images from a prompt list
import os
from huggingface_hub import InferenceClient

HF_TOKEN = os.environ.get("HF_TOKEN", "hf_xMB…bGTA")
client = InferenceClient(provider="auto", api_key=HF_TOKEN)

prompts = [
    ("A serene Japanese garden at dawn", "garden.png"),
    ("Futuristic Mars colony, domed cities", "mars.png"),
    ("Steampunk airship over Victorian London", "airship.png"),
]

for prompt, filename in prompts:
    print(f"Generating: {prompt}")
    image = client.text_to_image(prompt, model="Qwen/Qwen-Image")
    image.save(filename)
    print(f"  β†’ Saved {filename}")

print("Done!")
Note: HuggingFace GPU inference has rate limits. For high-volume generation, consider a dedicated Space with your own GPU.

Chapter 4 Β· Audio Processing

Audio Processing Pipeline

Speech-to-text via Whisper Large V3 Turbo and text-to-speech via MMS TTS + SpeechT5. All served through HuggingFace's GPU-accelerated inference.

Speech-to-Text: Whisper Large V3 Turbo

from huggingface_hub import InferenceClient

client = InferenceClient(api_key="hf_xMB…bGTA")

# Transcribe an audio file
result = client.automatic_speech_recognition(
    "meeting-recording.mp3",
    model="openai/whisper-large-v3-turbo"
)
print(result["text"])

# Supported formats: mp3, wav, flac, ogg, m4a, webm
# Languages: 99 languages including EN, ZH, JA, KO, ES, FR, DE, AR, HI
# Accuracy: ~4% WER on English (near-human)
CapabilityWhisper V3 Turbo
Languages99
English WER~4%
Max file size25 MB (HF limit)
TimestampsWord-level (via API)
TranslationAny β†’ English (built-in)

Text-to-Speech: MMS + SpeechT5

MMS TTS (English, simple API)

# Simple English TTS
audio = client.text_to_speech(
    "Hello! I'm CUTEADMOA-5.2, your multi-model AI assistant.",
    model="facebook/mms-tts-eng"
)
with open("intro.wav", "wb") as f:\n    f.write(audio)

SpeechT5 (multilingual, speaker embeddings)

# Multilingual TTS with voice cloning
from transformers import SpeechT5Processor, SpeechT5HifiGan, SpeechT5ForTextToSpeech
import torch

processor = SpeechT5Processor.from_pretrained("microsoft/speecht5_tts")
model = SpeechT5ForTextToSpeech.from_pretrained("microsoft/speecht5_tts")
vocoder = SpeechT5HifiGan.from_pretrained("microsoft/speecht5_hifigan")

# Load speaker embedding for voice character
speaker_embeddings = torch.load("speaker_embeddings.pt")

inputs = processor(text="Bonjour le monde!", return_tensors="pt")
speech = model.generate_speech(inputs["input_ids"], speaker_embeddings, vocoder=vocoder)

import soundfile as sf
sf.write("french_output.wav", speech.numpy(), 16000)
FeatureMMS TTSSpeechT5
LanguagesEnglish onlyMultilingual (FR, DE, ES, NL, +)
Voice cloningβ€”βœ“ Speaker embeddings
API simplicityVery simple (1 call)Moderate (needs embeddings)
QualityNatural ENNatural, expressive
HF Space neededMinimalRecommended for custom voices

End-to-End Audio Pipeline

Transcribe β†’ Process β†’ Generate response β†’ Speak. All on HuggingFace GPU.

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚ AUDIO β”‚ ───► β”‚ WHISPER β”‚ ───► β”‚ TEXT LLM β”‚ ───► β”‚ MMS TTS β”‚ ───► πŸ”Š VOICE β”‚ INPUT β”‚ β”‚ V3 Turbo β”‚ β”‚ (any 7) β”‚ β”‚ /SpeechT5β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ .mp3/.wav STT Reasoning TTS Output HF GPU (GPU) NVIDIA/HF/etc. HF GPU (GPU)
#!/usr/bin/env python3
# voice-assistant-pipeline.py
from huggingface_hub import InferenceClient
from openai import OpenAI

hf = InferenceClient(api_key="hf_xMB…bGTA")
nvidia = OpenAI(base_url="https://integrate.api.nvidia.com/v1", api_key="nvapi-…")

# 1. Transcribe
text = hf.automatic_speech_recognition("input.wav", model="openai/whisper-large-v3-turbo")["text"]

# 2. Process with LLM
response = nvidia.chat.completions.create(
    model="nvidia/nemotron-3-ultra-550b-a55b",
    messages=[{"role": "user", "content": text}]
)
reply = response.choices[0].message.content

# 3. Speak response
audio = hf.text_to_speech(reply, model="facebook/mms-tts-eng")
with open("output.wav", "wb") as f:\n    f.write(audio)\n\nprint(f"πŸŽ™οΈ Input: {text[:80]}...")
print(f"πŸ€– Reply: {reply[:80]}...")
print("πŸ”Š Saved to output.wav")

Chapter 5 Β· Ecosystem Integration

Claude Code via ACP

Connect CUTEADMOA-5.2 to Claude Code via the Agent Communication Protocol. Use any of the agent's models as the backend for Claude Code coding sessions.

Local stdio Setup

# Start Claude Code in ACP mode
claude-code --acp-stdio

# From OpenClaw agent, spawn Claude Code sub-agent with DeepSeek V4
sessions_spawn(
  runtime: "acp",
  agentId: "claude",
  model: "custom__ffe02b89-8b42-4f23-b4e9-af3a1bd1bd3d/deepseek-ai/deepseek-v4-pro",
  task: "Refactor the authentication module to use JWT with refresh tokens",
  taskName: "auth-refactor"
)

Recommended Claude Code Models

TaskModelWhy
Complex refactorsDeepSeek V4 Pro1M context, strong reasoning
Full codebase reviewQwen3 Coder2M context, code-optimized
Architecture designNemotron 550BDeep analytical reasoning
Quick fixesClaude Opus 4.8Fast, high-quality

Hermes Evolution System

CUTEADMOA-5.2 supports Hermes self-evolution. The agent watches for repeated patterns, costly mistakes, and user preferences β€” then drafts structured proposals for human approval.

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” Detection β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” Review β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚ AGENT β”‚ ────────────► β”‚ PROPOSAL β”‚ ────────► β”‚ HUMAN β”‚ β”‚ detects β”‚ β”‚ DRAFT β”‚ β”‚ approves β”‚ β”‚ pattern β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚ β”Œβ”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β” β”‚ APPLY β”‚ β”‚ Atomic β”‚ β”‚ Git-safe β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
Current intensity: cautious (50%) β€” proposals only after clear patterns emerge.

MCP Server Integration

CUTEADMOA-5.2 discovers and uses MCP servers for filesystem access, GitHub operations, database queries, web search, and more β€” all through a standardized JSON-RPC 2.0 protocol.

# openclaw.json β€” MCP configuration
{
  "mcp": {
    "enabled": true,
    "gatewayUrl": "http://localhost:3000/mcp",
    "servers": [
      {
        "name": "filesystem",
        "command": "npx",
        "args": ["@modelcontextprotocol/server-filesystem", "/workspace"]
      },
      {
        "name": "github",
        "command": "npx",
        "args": ["@modelcontextprotocol/server-github"],
        "env": { "GITHUB_TOKEN": "${GITHUB_TOKEN}" }
      },
      {
        "name": "postgres",
        "command": "npx",
        "args": ["@modelcontextprotocol/server-postgres"],
        "env": { "DATABASE_URL": "${DATABASE_URL}" }
      }
    ],
    "autoDiscover": true
  }
}

Filesystem

Read, write, list, glob, grep files in workspace.

@modelcontextprotocol/server-filesystem

GitHub

Issues, PRs, repos, files via GitHub API.

@modelcontextprotocol/server-github

Databases

PostgreSQL and SQLite query servers.

@modelcontextprotocol/server-postgres

Web

Brave Search and HTTP Fetch servers for internet access.

@modelcontextprotocol/server-brave-search

Direct API Access

Every provider can be called via OpenAI-compatible API. Provider URLs and keys are in the agent config.

# NVIDIA (Nemotron, DeepSeek, GLM 5.2)
from openai import OpenAI
client = OpenAI(base_url="https://integrate.api.nvidia.com/v1", api_key="nvapi-…")
response = client.chat.completions.create(
    model="deepseek-ai/deepseek-v4-pro",
    messages=[{"role": "user", "content": "Explain async/await in Python"}]
)

# Alibaba Cloud (Qwen3 Coder, Qwen3.5 Omni)
client = OpenAI(
    base_url="https://ws-hflmx206bvspck9t.ap-southeast-1.maas.aliyuncs.com/compatible-mode/v1",
    api_key="sk-ws-…"
)
response = client.chat.completions.create(
    model="qwen3-coder-next",
    messages=[{"role": "user", "content": "Write a FastAPI CRUD app"}]
)

# OpenRouter (Claude Opus 4.8)
client = OpenAI(base_url="https://openrouter.ai/api/v1", api_key="sk-or-…")
response = client.chat.completions.create(
    model="anthropic/claude-opus-4.8-fast",
    messages=[{"role": "user", "content": "Analyze this legal document..."}]
)

Provider Endpoints

ProviderBase URLModels
NVIDIAhttps://integrate.api.nvidia.com/v1Nemotron, DeepSeek, GLM 5.2
Alibaba Cloudhttps://ws-hflmx206bvspck9t.ap-southeast-1.maas.aliyuncs.com/compatible-mode/v1Qwen3 Coder, Qwen3.5 Omni
OpenRouterhttps://openrouter.ai/api/v1Claude Opus 4.8, Free tier
HuggingFace Imageshttps://router.huggingface.co/hf-inferenceQwen-Image, FLUX, SD 3.5
HuggingFace Audiohttps://router.huggingface.co/hf-inferenceWhisper, MMS, SpeechT5
Z.AI Autohttps://autoglm-api.autoglm.ai/autoclaw-proxy/proxy/autoclawzai_auto, GLM-5-Turbo, GLM-5.2
OpenAI Codexhttps://chatgpt.com/backend-apiGPT-5.5, GPT-5.4-Mini

Chapter 6 Β· Deployment

Cloud & Docker Deployment

CUTEADMOA-5.2 runs as an agent inside the AutoClaw gateway. Deploy the gateway, and all registered agents (including 5.2) are available.

# docker-compose.yml
services:
  gateway:
    image: openclaw/gateway:latest
    ports:
      - "8080:8080"
      - "8081:8081"  # ACP WebSocket
    volumes:
      - ./openclaw.json:/config/openclaw.json
      - ./agents/cuteadmoa-5-2:/agents/cuteadmoa-5-2
      - workspace-data:/workspace
    environment:
      - GATEWAY_MODE=production
      - NODE_ENV=production
    restart: unless-stopped

  nginx:
    image: nginx:alpine
    ports:
      - "443:443"
      - "80:80"
    volumes:
      - ./nginx.conf:/etc/nginx/nginx.conf
      - ./certs:/etc/nginx/certs
    depends_on: [gateway]

volumes:
  workspace-data:

HuggingFace Space (GPU)

For image/audio models needing dedicated GPU, create a HuggingFace Space. Copy the model configs and set the token as a Space secret. This gives you dedicated GPU inference without rate limits.

# Create a HuggingFace Space with GPU
# 1. Go to https://huggingface.co/new-space
# 2. Choose "Docker" SDK, "T4 small" or "A10G" hardware
# 3. Set HF_TOKEN as a Space secret
# 4. Deploy your inference endpoint

# Space Dockerfile
FROM huggingface/transformers-pytorch-gpu:latest
RUN pip install huggingface_hub Pillow soundfile
COPY app.py /app/app.py
CMD ["python", "/app/app.py"]

# Then update models.json baseUrl to your Space URL
# "baseUrl": "https://your-username-your-space.hf.space"

Security & Token Management

⚠️ Tokens are sensitive. The auth-profiles.json file contains API keys. Never commit it to version control. The keys below are truncated for documentation.

Token Inventory

ProviderKey PrefixRotationNotes
NVIDIAnvapi-…WZMNEvery 90 daysShared for Nemotron, DeepSeek, GLM
NVIDIA (GLM)nvapi-…DEfMEvery 90 daysDedicated for GLM 5.2
OpenRoutersk-or-…6440Every 60 daysClaude Opus + Free tier
Alibaba Cloudsk-ws-…hSIwPer projectQwen models
HuggingFacehf_xMB…bGTAEvery 180 daysAll image + audio models

Best Practices


Appendix A Β· API Reference

File Structure

~/.openclaw-autoclaw/agents/cuteadmoa-5-2/
β”œβ”€β”€ agent/
β”‚   β”œβ”€β”€ auth-profiles.json    # All API keys (encrypted references)
β”‚   β”œβ”€β”€ models.json           # All 15 model definitions
β”‚   └── plugins/
β”‚       └── zai/
β”‚           └── catalog.json   # Z.AI auto-routing catalog
β”œβ”€β”€ workspace/
β”‚   β”œβ”€β”€ AGENTS.md             # Agent behavior rules
β”‚   β”œβ”€β”€ SOUL.md               # Personality and tone
β”‚   β”œβ”€β”€ IDENTITY.md           # Agent identity record
β”‚   β”œβ”€β”€ USER.md               # User profile
β”‚   β”œβ”€β”€ TOOLS.md              # Local tool configurations
β”‚   β”œβ”€β”€ HEARTBEAT.md          # Periodic task checklist
β”‚   └── docs/
β”‚       β”œβ”€β”€ CUTEADMOA-5.2-Documentation.html
β”‚       └── CUTEADMOA-5.2-Documentation.pdf
└── sessions/
    └── sessions.json          # Session state

Model IDs (full paths)

PurposeModel ID
Auto-routerzai/zai_auto
Nemotron 550Bcustom__0716bfa4/nvidia/nemotron-3-ultra-550b-a55b
GLM 5.2custom__c553fcc0/z-ai/glm-5.2
DeepSeek V4 Procustom__ffe02b89/deepseek-ai/deepseek-v4-pro
Qwen3 Codercustom__857372ba/qwen3-coder-next
Qwen3.5 Omni+custom__2cda56e4/qwen3.5-omni-plus
Claude Opus 4.8custom__b66c850c/anthropic/claude-opus-4.8-fast
OpenRouter Freecustom__fb75a144/openrouter/free
GPT-5.5codex/gpt-5.5
GPT-5.4-Minicodex/gpt-5.4-mini
Qwen-Imagehuggingface-images/Qwen/Qwen-Image
FLUX.1-devhuggingface-images/black-forest-labs/FLUX.1-dev
SD 3.5 Largehuggingface-images/stabilityai/stable-diffusion-3.5-large
Whisper V3 Turbohuggingface-audio/openai/whisper-large-v3-turbo
MMS TTShuggingface-audio/facebook/mms-tts-eng
SpeechT5 TTShuggingface-audio/microsoft/speecht5_tts

Appendix B Β· Version Compatibility Matrix

ComponentMin VersionTestedNotes
AutoClaw Desktop2026.62026.7ACP + Hermes since 2025.3
Claude Code1.1.01.2.15ACP WebSocket in 1.1.0
Node.js2022.22LTS recommended
Python3.103.123.9 breaks MCP SDK
Docker24.027.xBuildKit required
macOS1325.5 (Arm64)Apple Silicon native
HuggingFace Hub0.200.28+pip install huggingface_hub
MCP Spec2024-11-052025-06-18Backward compatible

Appendix C Β· Troubleshooting

IssueCauseFix
Agent not in listNot registered in openclaw.jsonVerify agents.list has cuteadmoa-5-2
Model not foundmodels.json path mismatchCheck provider UUID matches config
HuggingFace 429Rate limit hitCreate a dedicated HF Space with GPU
DeepSeek V4 failsNVIDIA API key issueVerify key in auth-profiles.json
Image blank/errorPrompt NSFW filter or model timeoutTry a different model, simplify prompt
Whisper no outputAudio format unsupportedConvert to .mp3 or .wav (16kHz mono)
TTS garbled audioWrong model parametersUse MMS for simple EN, SpeechT5 for multi
ACP connection refusedClaude Code not in PATHnpm install -g @anthropic-ai/claude-code
Session hangsSub-agent timeoutIncrease runTimeoutSeconds in config

Debug Commands

# Verify agent registration
openclaw agents list | grep cuteadmoa

# Test a model directly
curl -s https://integrate.api.nvidia.com/v1/models \
  -H "Authorization: Bearer nvapi-YOUR_KEY" | python3 -m json.tool

# Test HuggingFace image
python3 -c "
from huggingface_hub import InferenceClient
c = InferenceClient(api_key='hf_xMB…bGTA')
img = c.text_to_image('test red circle', model='Qwen/Qwen-Image')
img.save('/tmp/test.png')
print('OK')
"

# Check gateway logs
tail -f ~/.openclaw-autoclaw/logs/gateway.log | grep -i "cuteadmoa\|error\|acp"

CUTEADMOA-5.2 Β· v5.2 Β· Generated 2026-07-29 Β· Multi-Model AI Agent

15 models Β· 9 providers Β· Text Β· Image Β· Audio Β· Default routing: zai/zai_auto

⬀ Active