Blackbox AI

New

New · Kimi K2.6 is now available on Blackbox AI. Frontier reasoning, one API.

AGENT-01 // REFACTOR

blackbox agent run --task refactor-auth

[INIT] Loading codebase context...

[SCAN] 47 files analyzed in 1.2s

[PLAN] Extracting auth middleware → /lib/auth.ts

[EDIT] src/routes/api.ts — removed inline checks

[EDIT] src/middleware/auth.ts — created guard

[TEST] 12/12 passing

[DONE] Refactor complete. PR ready.

AGENT-02 // MIGRATE

blackbox agent run --task db-migration

[INIT] Connecting to schema registry...

[DIFF] 3 tables modified, 1 added

[GEN] Migration 0047_add_teams.sql

[VALIDATE] Foreign keys ............. OK

[VALIDATE] Indexes .................. OK

[APPLY] Dry run successful

[DONE] Migration staged.

AGENT-03 // TEST-GEN

blackbox agent run --task generate-tests

[SCAN] Uncovered functions: 23

[GEN] tests/auth.test.ts (8 cases)

[GEN] tests/billing.test.ts (6 cases)

[GEN] tests/api.test.ts (9 cases)

[RUN] 23/23 passing

[COV] Coverage: 47% → 89%

[DONE] Test suite committed.

AGENT-04 // DEPLOY

blackbox agent run --task deploy-staging

[BUILD] next build .................. OK

[LINT] 0 errors, 0 warnings

[TYPE] tsc --noEmit ................ OK

[PUSH] Deploying to staging...

[DNS] https://staging.blackbox.dev

[HEALTH] 200 OK — 43ms

[DONE] Live on staging.

CHAIRMAN LLM // JUDGE

chairman evaluate --round 1

[RECV] 4 agent submissions received

[EVAL] Agent-01: refactor .......... 9.2/10

[EVAL] Agent-02: migration ......... 8.8/10

[EVAL] Agent-03: test-gen .......... 9.5/10

[EVAL] Agent-04: deploy ............ 9.1/10

[RANK] Best: Agent-03 (test coverage)

[VERDICT] All agents passed. Merging.

SYSTEM // MONITOR

blackbox monitor --live

[SYS] CPU: 12% | MEM: 3.2GB / 16GB

[SYS] Active agents: 4

[SYS] Queue depth: 0

[NET] API latency p99: 89ms

[NET] Requests/min: 2,847

[COST] Session: $0.42

[STATUS] All systems nominal.

AGENT-05 // REVIEW

blackbox agent run --task code-review

[LOAD] PR #247 — 14 files changed

[SCAN] Security patterns ........... OK

[SCAN] Performance anti-patterns ... 1 found

[WARN] Avoid N+1 in /api/teams.ts:34

[SCAN] Type coverage ............... 100%

[APPROVE] No blockers found

[DONE] Review posted.

AGENT-06 // DOCS

blackbox agent run --task update-docs

[SCAN] 7 undocumented exports found

[GEN] docs/api-reference.md updated

[GEN] docs/auth-guide.md created

[GEN] README.md — added quickstart

[LINK] Cross-references validated

[SPELL] 0 issues

[DONE] Documentation shipped.

AGENT-07 // SECURITY

blackbox agent run --task security-audit

[SCAN] Dependencies: 847 packages

[CVE] 0 critical, 0 high, 2 low

[WARN] lodash@4.17.20 — prototype pollution

[AUTH] Token rotation .............. OK

[AUTH] CORS policy ................. OK

[SECRETS] No exposed credentials

[DONE] Audit passed with warnings.

AGENT-08 // PERF

blackbox agent run --task perf-optimize

[PROFILE] Lighthouse run complete

[SCORE] Performance: 94 → 99

[FIX] Lazy-loaded 3 heavy components

[FIX] Replaced unoptimized images

[FIX] Tree-shook 47KB dead code

[BUNDLE] 312KB → 198KB (-37%)

[DONE] Performance optimized.

AGENT-09 // SCAFFOLD

blackbox agent run --task scaffold-service

[INIT] Template: microservice-ts

[GEN] src/index.ts — entry point

[GEN] src/routes/ — REST handlers

[GEN] src/db/ — Drizzle schema

[GEN] Dockerfile + compose.yml

[GEN] .github/workflows/ci.yml

[DONE] Service scaffolded.

AGENT-10 // TRANSLATE

blackbox agent run --task i18n-extract

[SCAN] 134 hardcoded strings found

[GEN] locales/en.json (134 keys)

[GEN] locales/fr.json (auto-translated)

[GEN] locales/de.json (auto-translated)

[WRAP] Components updated with t()

[VALIDATE] No missing keys

[DONE] i18n extraction complete.

AGENT-11 // ROLLBACK

blackbox agent run --task rollback-prod

[ERR] deploy-v2.42.0 health check failed

[FETCH] Last stable: deploy-v2.41.0

[DIFF] 3 commits to revert

[REVERT] Applied reverse patches

[BUILD] Rebuild ..................... OK

[PUSH] Deploying rollback...

[DONE] Rolled back to v2.41.0.

AGENT-12 // LINT-FIX

blackbox agent run --task auto-lint

[SCAN] 2,847 files checked

[FIX] 47 auto-fixable issues

[FIX] Removed unused imports (23)

[FIX] Consistent quote style (18)

[FIX] Trailing commas (6)

[RUN] eslint --fix ................ OK

[DONE] Codebase clean.

AGENT-13 // CANARY

blackbox agent run --task canary-release

[BUILD] Production build ........... OK

[SPLIT] 5% traffic → canary

[WATCH] Error rate: 0.00%

[WATCH] p99 latency: 47ms

[PROMOTE] Canary → 100%

[CLEANUP] Old revision removed

[DONE] Canary promoted.

AGENT-14 // SCHEMA

blackbox agent run --task validate-schema

[LOAD] OpenAPI spec v3.1

[CHECK] Endpoints: 34 documented

[CHECK] Types match runtime: 34/34

[CHECK] Breaking changes: 0

[GEN] SDK types regenerated

[DIFF] No client changes needed

[DONE] Schema validated.

SYSTEM // HEARTBEAT

blackbox heartbeat --interval 1s

[PING] api.blackbox.dev ............ 8ms

[PING] auth.blackbox.dev ........... 11ms

[PING] cdn.blackbox.dev ............ 4ms

[PING] ws.blackbox.dev ............. 9ms

[UPTIME] 99.997% (30d rolling)

[CERTS] All valid > 90 days

[STATUS] All services alive.


Your Agent Platform

Run agents from anywhere, anytime, autonomously.

One platform, six surfaces. Dispatch autonomous coding agents from your terminal, IDE, cloud, API, phone, or browser. They compete, collaborate, and ship code — while you focus on what matters.

CLI

Your terminal, supercharged

Dispatch competing agents from a single command. They analyze your codebase, generate solutions in parallel, and open PRs — no browser needed.

  • Multi-agent parallel execution
  • Automatic PR creation
  • CI/CD pipeline integration

IDE

Agents in your editor

Agents work alongside you inside VS Code or Blackbox IDE. Real-time code generation, refactoring, and testing — right where you write code.

  • Inline code generation
  • Context-aware refactoring
  • Integrated test runner

Cloud

Always-on, always working

Deploy autonomous agents to the cloud. They monitor, fix, and optimize your codebase 24/7 — even while your team sleeps.

  • 24/7 autonomous operation
  • Automated monitoring & fixes
  • Team dashboards & controls

API

Programmable agent execution

Integrate agent execution into any workflow with OpenAI-compatible endpoints. Chat completions, multi-agent orchestration, and real-time streaming.

  • OpenAI-compatible endpoints
  • Multi-agent orchestration
  • WebSocket streaming

Mobile

Ship code from your pocket

Review agent work, approve PRs, and dispatch new tasks from anywhere. The Blackbox mobile app keeps you in control on the go.

  • Push notification alerts
  • PR review & approval
  • One-tap agent dispatch

Builder

Describe it, agents build it

Go from natural language to a deployed application. Agents handle the architecture, code, testing, and deployment — you describe the vision.

  • Natural language to app
  • Full-stack generation
  • One-click deployment

Chairman LLM

Run Agents in Parallel.

Dispatch the same task to multiple AI agents, then let Chairman LLM evaluate every candidate on correctness, performance, risk, and complexity. Best output wins.

Task

Implement rate limiting middleware with Redis backend for the API gateway

Solutions

Claude Code: I'll use a sliding window algorithm with Redis MULTI/EXEC for atomicity. The middleware checks req count per IP in a 60s window, returns 429 when exceeded.

Codex: Implementing token bucket via Redis INCR + EXPIRE. Each request decrements the bucket; refill rate is configurable per route. Includes retry-after header.

Blackbox: I recommend a distributed rate limiter using Redis sorted sets for precise sliding windows. Supports per-user and per-endpoint limits with graceful degradation.

Winner Selected

claude code

confidence: 0.94

TESTS: 46/46 correctness: 0.97

Evaluating Solutions

  • Claude Code: 0.94
  • Codex: 0.81
  • Blackbox: 0.78