DeepSeek Harness: Architecture, Tool Execution, and Setup Guide
Comprehensive technical guide to DeepSeek Harness (dsh): Cordis micro-kernel architecture, append-only trajectory tracing, pluggable model adapters, and local deployment.


Synthesizing article benchmarks & model metrics...
DeepSeek has open-sourced DeepSeek Harness (dsh), a modular, local-first agent runtime engineered from the ground up on one core premise: everything in an autonomous agent is a swappable plugin.
The project landed on GitHub as an MIT-licensed developer preview (deepseek-ai/deepseek-harness) and rapidly accumulated over 140,000 GitHub stars. While proprietary coding assistants like Claude Code, Cursor, and Windsurf operate as tightly coupled commercial monoliths, DeepSeek Harness provides an open, hackable foundation where memory layers, tool registries, sandboxes, and foundation models can be replaced with simple configuration files.
DeepSeek frames the core philosophy around an elegant equation:
$$\text{Agent} = \text{Model} + \text{Harness}$$
A model by itself is simply a statistical next-token predictor. To create an agent, you require the harness: the environment interaction, file system access, terminal execution, error recovery loops, and deterministic trajectory tracing that allow the model to interact with software systems. Explore where DeepSeek models rank on our AI Model Leaderboard and check individual scorecards for DeepSeek V4 Pro 0813 and DeepSeek V4 Flash.
graph TD
A[DeepSeek Harness Runtime] --> B[Cordis Kernel]
B --> C[Model Plugins: DeepSeek / Claude / GPT / Local]
B --> D[Tool Plugins: Shell / File Search / Git]
B --> E[Memory & Context Plugins: Local Traces / KV Cache]
B --> F[Sandbox Plugins: Docker / MicroVM / Local]
B --> G[UI Plugins: Web Dashboard / CLI / Headless]
Architectural Deep-Dive: The Cordis Micro-Kernel
Under the hood, DeepSeek Harness runs on Cordis, a lightweight TypeScript micro-kernel designed around spatiotemporal composability. Rather than rigid class inheritance or complex dependency injection frameworks, Cordis delivers:
- Dynamic Plugin Mounting: Mount and unmount tools, memory modules, and model providers at runtime without restarting the agent session.
- Service Discovery: Plugins register services (such as
ctx.storage,ctx.sandbox, andctx.models) that other plugins can seamlessly consume. - Event Lifecycle Interception: Intercept agent thoughts, tool execution results, and error states using structured event hooks.
+-------------------------+ | DeepSeek Harness | +-------------------------+ | +-------------------------+ | Cordis Kernel | +-------------------------+ | +------------+-----------+-----------+------------+ | | | | | [Models] [Tools] [Skills] [Memory] [Sandboxes] (DeepSeek/V4) (Shell/Git) (Subagents) (Session) (Container)Full Trajectory Tracing: Deterministic Step Replay
Debugging multi-step coding agents is notoriously difficult when systems operate as black boxes. DeepSeek Harness solves this with first-class trajectory tracing. Every single event during an agent run is recorded in an immutable, append-only JSONL session log:
- Exact System Prompts & Context: Ingested repository files and prompt token counts.
- Model Thinking Chains: Full internal reasoning traces prior to tool calls.
- Tool Invocations & Exit Codes: Exact bash arguments, stdout, stderr, and execution duration.
- Subagent Delegation Traces: Hierarchy of spawned sub-tasks and return values.
The Power of Step Forking
Because trajectories are deterministic, developers can:
- Inspect the complete decision graph in real time through the local Web UI (
http://127.0.0.1:3080). - Fork an active trajectory at step 8, adjust the system prompt or tool parameter, and re-execute to test alternate problem-solving paths.
- Replay regression benchmarks across model updates to verify consistency.
The 4 Built-In Runtime Modes
DeepSeek Harness ships with four distinct profiles:
| Runtime Mode | Primary Use Case | Enabled Capabilities |
|---|---|---|
| Standard Mode | Day-to-day software engineering | Full agent loop: bash execution, multi-file editing, AST search, subagents. |
| Code Mode | Programmatic automation | Exposes tools via a TypeScript SDK, enabling models to generate executable scripts. |
| Minimal Mode | Model benchmarking | Barebones environment with shell and str_replace_editor for clean SWE evaluations. |
| Creator Mode | Agent development | Runtime inspector and sandbox for authoring custom plugins, presets, and workflows. |
Multi-Model Agnosticism: Beyond DeepSeek
While engineered to extract maximum performance from DeepSeek V4 Pro 0813 (74.7% DeepSWE) and DeepSeek V4.1 Flash, DeepSeek Harness is strictly model-agnostic.
Through its modular provider layer, developers can route tasks to:
- DeepSeek API:
deepseek-v4-pro,deepseek-v4-flash - Anthropic Claude:
claude-sonnet-5,claude-fable-5 - OpenAI / OpenRouter:
gpt-5-6-sol,o3-mini - Local Self-Hosted Endpoints: vLLM, SGLang, Ollama, LM Studio
Quickstart: Running DeepSeek Harness in Under 60 Seconds
Method 1: Instant Local Web UI (Zero Install)
Requires Node.js (v18+) installed:
npx @deepseek-ai/dsh webOpen your browser to http://127.0.0.1:3080 to access the interactive visual agent dashboard.
Method 2: Development Installation from Source
To author custom plugins or modify the agent loop:
git clone https://github.com/deepseek-ai/deepseek-harness.gitcd deepseek-harnesspnpm installpnpm run buildpnpm dsh webFinal Verdict
DeepSeek Harness establishes the open-source standard for autonomous agent scaffolding.
By untangling model weights from the execution runtime and packaging capabilities behind the Cordis micro-kernel, DeepSeek provides developers with a production-ready, MIT-licensed foundation to build next-generation software agents.
Compare model benchmarks in our AI Model Leaderboard and read our head-to-head comparison of ZCode vs DeepSeek Harness.
Benchmark and pricing data is aggregated from OpenRouter, Artificial Analysis, and models.dev, then scored with the published RankLLMs Index. These are the primary sources behind the numbers in this article.
- •DeepSeek Harness Official GitHub Repository(Primary Source →)
- •DeepSeek Harness Developer Documentation(Primary Source →)
- •Cordis Micro-Kernel Framework(Primary Source →)

Lucky Yaduvanshi(luckyyaduvanshi.in →)
Founder of RankLLMs • AI Researcher & Software Engineer focusing on LLM benchmarking, DevOps, and autonomous coding agents.
Recommended Reading
Continue exploring related technical benchmarks, model deep-dives, and autonomous coding agent guides.

ZCode vs DeepSeek Harness: Integrated IDE vs Modular Agent Framework
ZCode and DeepSeek Harness represent two contrasting futures for AI coding agents: a full-stack desktop ADE vs a modular, plugin-based Cordis runtime. Here is how they compare in architecture, benchmarks, and real-world developer workflows.
Lucky Yaduvanshi
DeepSeek V4 Pro 0813 vs GLM-5.3: Frontier Agent Benchmarks and Architecture
Frontier agent comparison: DeepSeek V4 Pro 0813 (1.6T MoE) vs Zhipu AI's GLM-5.3 across Terminal-Bench, SWE-bench, reasoning, and API economics.
Lucky Yaduvanshi
GLM-5.3 in ZCode: Agentic Coding Integration and Evaluation
Z.ai rolled out GLM-5.3 to all ZCode users with free tier access, reset quotas, and top scores on CyberGym (84.5%), GDPval-AA, and Terminal-Bench 2.1 (88.2).
Lucky Yaduvanshi
DeepSeek V4 Pro 0813: Architectural Analysis, Benchmarks, and $0.435/1M Economics
DeepSeek V4 Pro 0813 achieves 87.9 on Terminal-Bench 2.1 using a 1.6T MoE architecture at $0.435/1M tokens. Here is the technical report and benchmark analysis.
Lucky Yaduvanshi