My Claude Code review, after running it on this site's research, writing, browser checks and data scripts: it does real multi-agent work well, and it states wrong things with full confidence, so nothing it produces ships here until an independent check has passed it. With those checks built in, it is worth using; without them, you become the check. This is a review of Claude Code as the agent behind one person's marketing analytics site, not of a large codebase, and every product claim links to the official docs as read on October 3, 2026. If you were looking for Code Review, the pull request reviewer, it has its own section below.
What is Claude Code, and what do I use it for?
Claude Code is Anthropic's agentic coding tool: it reads your codebase, edits files, runs commands and works with your development tools (overview). The docs add that while it excels at coding, it can help with anything you can do from the command line, including writing docs and researching topics (how Claude Code works). That is most of my use. It writes and fact-checks this blog, renders pages in a browser, runs read-only tag audits of public pages for outreach research, and runs plain scripts against real APIs.
Is it good at research and writing, not just code?
Yes, when the work is split across agents. Each subagent runs in its own context window with its own system prompt and tool access (subagents), and the docs' dynamic workflows run dozens to hundreds of agents from one script (workflows). Every article here gets one writer agent, an independent fact-checker that re-reads the official docs and fixes or cuts claims, and editor passes. One run on October 2, 2026 used 32 agents to write 13 manuals, tighten 5 drafts and render-check 20 pages.
Can it check pages in a real browser?
Yes. In that workflow one agent renders every page in a real browser at desktop and phone width. For outreach research, agents run read-only tag audits of public pages in headless Chrome with network capture, and they never submit a form. The docs count a browser screenshot among the checks Claude can read, next to a test suite, a build exit code and a linter (best practices). If you want it in your own browser instead, the Claude in Chrome extension connects it to Chrome and shares your login state (Chrome).
Can it work with real APIs and data?
Yes, through plain scripts. Claude Code's execution tools run shell commands, start servers, run tests and use git (tools). My data work runs that way: the Search Console API with a service account, Bing URL Submission, and a keyword data API behind a cost gate that refuses any run over a limit unless I approve it. Third-party SEO skills run in this project too, under a rules block of mine: no paid API call without approval, and my positioning rules win over the skill's.
Where did Claude Code fail without a check?
Each of these read fine to the agent that produced it:
- Confidently wrong. A Google Ads claim about which calls count missed Google's newer AI rating of recorded calls (AI-qualified call leads). A fact-checker caught it.
- True only from where it ran. A finding that a site sent consent "denied" and showed no banner held only because the check ran from Bulgaria: the Wix site set denied for EU visitors only. Wix's help says that without a cookie banner, a visitor from a country that requires consent sends no data to Google Analytics (Wix help). A fact-checker caught that one too.
- Padded. I rejected early drafts as padded. The fix became a fixed manual format and a saved rule.
- Invented frequency. A sentence claimed "the case I see most" with no count behind it. A fact-checker cut it.
The docs describe the root cause plainly: Claude stops when the work looks done, and without a check it can run, "looks done" is the only signal available (best practices). Their fix for plausible-looking work that misses edge cases is just as blunt: "If you can't verify it, don't ship it" (common failure patterns).
What did I put around it?
Three layers:
- Verifiers. A fact-checker per article works from a fresh context, so the agent that wrote a claim is not the one grading it. The docs recommend the same second opinion: a verification subagent that has a fresh model try to refute the result (best practices).
- A linter.
scripts/content_lint.pyfails the run on em dashes, links to removed or missing pages, over-long titles and descriptions, unknown categories, repeated FAQ questions, broken manual structure and banned phrasing, the invented-frequency phrase included. A dated baseline holds known errors in frozen pages. - Approvals. I approve every commit and push, and agents never push on their own.
Where do the rules live, and are they enforced?
My personal CLAUDE.md holds working rules: the language I talk in, never commit or push without explicit approval, no AI attribution in commits. The project CLAUDE.md holds positioning and copy rules and pulls in AGENTS.md with an @AGENTS.md line, the docs' @path/to/import syntax (imports). Auto memory adds a folder with one file per memory and a MEMORY.md index whose first 200 lines or 25KB load into every session (auto memory). Mine holds rules such as running the dev server on port 3100 and never 3000, requesting recrawls after every deploy, checking consent from a non-EU location and writing manuals without filler. The docs are clear about the limit: Claude treats both "as context, not enforced configuration" (memory), while permission rules are enforced by Claude Code, not by the model (permissions). So the rules that must hold also live in the linter and in my approvals, not only in files Claude reads.
Is this the same as Claude Code's Code Review?
No. Code Review is a managed service that analyzes GitHub pull requests and posts findings as inline comments, in research preview for Team and Enterprise subscriptions (Code Review). Multiple agents look for different classes of issue in parallel, then a verification step checks candidates against actual code behavior to filter out false positives; findings are tagged by severity and do not approve or block the PR (how reviews work). On other plans, /code-review reviews a diff in your terminal as a background subagent, and --fix applies the findings to your working tree (review a diff locally). This review does not cover it, but its shape matches my article pipeline: one step finds or writes, a separate step tries to prove it wrong.
Is Claude Code worth it?
For this site, yes, on one condition: every output is a draft until something other than the agent that made it has checked it. The setup that earns that trust is small: a writer, a verifier with a fresh context, a script that fails the run, and a person who approves the push. How the writer and the verifier split the work is in Claude Code subagents.
Tags
Frequently asked questions
Is Claude Code only for coding?
No. Anthropic's docs say that while it excels at coding, it can help with anything you can do from the command line, such as writing docs, running builds, searching files and researching topics. I use it for research, for writing and fact-checking articles, for browser render checks and for scripts against the Search Console and Bing APIs.
Does Claude Code make things up?
It can state wrong things with full confidence. On this blog one draft missed Google's newer AI rating of recorded calls, and another invented a frequency with no count behind it. The docs' answer to plausible-looking work is a check Claude can run, such as tests, a linter or a verification subagent, and not shipping what you cannot verify.
Is Claude Code the same as Claude Code Review?
No. Claude Code is the agentic tool that reads, edits and runs commands in your project. Code Review is a managed service in research preview for Team and Enterprise subscriptions that reviews GitHub pull requests with multiple agents and posts inline comments. On other plans, the /code-review command reviews a diff locally.
Can Claude Code push to GitHub without asking me?
In Manual mode it asks before any Bash command outside a built-in set of read-only commands, so a git push prompts you. Other modes and saved approvals can remove that prompt; an ask rule such as Bash(git push *) forces it back, even in auto mode. The docs warn that a push written another way, such as git -C . push, is not matched.