

- 13
- Episodes
- Daily
- Cadence
- 2026
- First episode
About Passing CCAR-F
An independent, unofficial study companion for the Claude Certified Architect, Foundations exam. Not affiliated with, sponsored by, or endorsed by Anthropic. brodynetworks.substack.com (https://brodynetworks.substack.com/s/passing-ccar-f?utm_medium=podcast)
- Publisher
- Brody Networks
- Category
- education · technology
- Language
- en
- Explicit
- No
- First episode
- 4 Sept 2026
- Latest episode
- 4 Sept 2026
Latest episodes
13 episodes in the feed.

4 Sept 2026
Episode 12. Exam Day, and Working the Questions
The finale. The exam-day experience, pacing sixty items across a hundred and twenty minutes, the five patterns that decide most scenario questions, and three walkthroughs built from the exam guide’s own publicly released sample questions, the only pre-existing items this series ever discusses. Independent and unofficial. This series is not affiliated with, sponsored by, or endorsed by Anthropic. Nothing in it is exam content. This finale discusses the exam guide’s own published sample questions from guide section 9, which Anthropic released publicly for exactly this purpose; every other episode’s practice questions were written for this show against the published guide, which is a free public document linked below. Chapters * 0:00 Cold open and disclaimer * 0:23 Exam day, the last episode * 0:54 The confidentiality agreement * 2:16 Rules of conduct and workspace * 2:50 Pacing: sixty items, two minutes each * 3:41 The five recurring patterns * 3:52 Pattern 1: deterministic beats probabilistic * 4:29 Pattern 2: fix the cause at its own layer * 5:04 Pattern 3: the proportionate first step * 5:44 Pattern 4: least privilege * 6:16 Pattern 5: structured data over prose * 6:44 Three walkthroughs from the guide’s own questions * 7:10 Walkthrough 1: identity verification, domain 1 * 8:12 Walkthrough 2: scattered test files, domain 3 * 9:11 Walkthrough 3: batch versus blocking, domain 4 * 10:10 Why the same five patterns keep appearing * 10:37 Retakes, renewal, and criterion-referenced scoring * 11:40 The close Sources * Claude Certified Architect, Foundations Exam Guide, Version 1.0 (section 9, published sample questions; sections 10, 12, 13, 14, 15) (https://everpath-course-content.s3-accelerate.amazonaws.com/instructor%2F6nizmqk8tpzpfjvt6qmmav7rh%2Fpublic%2F1783542750%2FClaude+Certified+Architect+%E2%80%93+Foundations+Exam+Guide.pdf) Transcript This is Passing C C A R F, an independent study companion for the Claude Certified Architect, Foundations exam. Sixty items, a hundred and twenty minutes, and a published blueprint that tells you almost exactly what it is going to ask. Independent and unofficial. Not affiliated with or endorsed by Anthropic. No exam content. Eleven episodes ago, this show was a hunch that a foundations exam could be studied for out loud. Today it is exam day. The last episode. The only one where we talk about actual questions. These particular ones were published by Anthropic itself, in the open, for exactly this purpose. No task statements today. This episode is the room, the clock, and the patterns that decide most of what you will see once you sit down. Start with the room, since it is the part nobody studies for. You will accept a confidentiality agreement before the exam begins, covering every question, answer option, and scenario as Anthropic’s confidential property. Decline it, and the session ends with no refund. Accept it, and from that moment forward, this show and every other honest source goes quiet on your specific paper, forever. You may be tempted, weeks later, to compare notes with someone else who has taken it. Do not. That conversation ends one of two credentials, and it is never worth the trade. That silence is not a policy inconvenience. It is why every question in this series was built from the published guide instead of a real one. And why it had to stay that way. There was never a version of this show that could wait until after someone sat the exam and still call itself independent. The scripting had to happen before, on the public record alone, or not at all. Not a rumor, not a forum post, not something somebody who sat the exam once passed along secondhand. Every fact in every one of these twelve episodes traces back to something Anthropic put on a public page. That is not a limitation this show worked around. It is the entire premise it was built on. Online or at a test center, the rules of conduct are the same in spirit. A clear workspace. No phones, no notes, no second monitor. Webcam in view the whole session if you are testing online. None of this is unusual for a proctored exam. Treat it as background noise, not as something to worry about the morning of. Every certification with real value runs this way. Familiar rules are not a warning sign. They are just the cost of a credential that means the same thing for everyone who holds it. Pacing is arithmetic, and it is worth doing once before you walk in. Sixty items, one hundred twenty minutes. Two minutes a item, on average, and averages are the operative word. Some items you will read once and know. Some will take four minutes of actually working through a scenario. Do not panic at minute forty because you have only finished fifteen items. Fifteen at minute forty is on pace, not behind, once you account for the reading. Panic itself costs more time than the pace ever did. The scenario-based items are front loaded with reading, and that reading is where the two minutes goes. Reading the scenario is not time lost before the real work starts. It is the real work. Most of what decides your answer is sitting in the setup, not in the four options underneath it. Now the five patterns. Not exam content. Just what eleven episodes of task statements add up to, once you stand back from all thirty of them at once. Pattern one: deterministic beats probabilistic, whenever consequences are financial or carry safety weight. A hook or a gate wins over a better sentence. Every single time the downside of a miss is real money or real harm. You heard this in episode three with the refund agent, and it is the most reliable single lens in the whole exam. Ask yourself one question before anything else: if this specific step gets skipped, does anyone lose money or get hurt. If the answer is yes, stop looking for the well written instruction. Look for the gate. Pattern two: fix the cause at its own layer, never compensate downstream. A narrow coordinator decomposition gets fixed at the coordinator, not by adding a fact-checking pass after synthesis. A missing tool description gets fixed by rewriting the description, not by adding a routing layer on top of it. Say the exam offers you a fix that sits one layer away from where the actual problem lives. That fix is almost always the trap. It is the answer that sounds responsible and touches nothing that actually broke. Pattern three: the proportionate first step. Cheap, high leverage fixes before new infrastructure. Expand a tool description before you build a classifier. Add explicit criteria before you deploy a second model to check the first one’s work. The exam consistently rewards the smallest change that actually reaches the root cause. It punishes reaching for machinery before you have tried the obvious thing. A second model, a new pipeline stage, a whole new subagent. All of these are real tools. All of them are wrong the moment a one line fix was sitting right there unused. Pattern four: least privilege, applied to tools and to context both. Scope each agent to what its role actually needs. Scope what a subagent receives to what it actually has to reason about. Every unnecessary tool is a temptation waiting to be misused. Every unnecessary block of context is competing for attention against the parts that matter. Scope is not caution for its own sake. It is removing a failure mode before it ever gets the chance to happen. Pattern five: structured data beats prose at every handoff. Between subagents, into a schema, across a summarization boundary, in an error response. Anywhere information crosses a seam, the version that survives is the one with fields, not the one trusted to a sentence. A well written sentence still has to be re-read and re-interpreted by whatever receives it. A field just has to be looked up. Hold those five, and you can work through unfamiliar scenarios by asking which one applies, rather than hunting for a memorized fact. Let’s prove it, on three of the guide’s own published sample questions. These are not questions written for this show. They are printed in the exam guide itself, publicly released. That is exactly why we can discuss them here, and nowhere else in the series. First, domain one. A support agent skips its identity verification step. It calls the order lookup using only a customer’s stated name, twelve percent of the time. It occasionally misidentifies accounts and issues refunds to the wrong person. The guide asks what change would most effectively fix this. The published answer is a programmatic prerequisite. It blocks the lookup and the refund calls until identity verification has actually returned a confirmed ID. A stronger system prompt and better few shot examples are both offered as tempting alternatives. Both are wrong for the same reason. They are still probabilistic, and this is a financial consequence problem. Pattern one, straight through the middle. Read that scenario again and the tell was there from the first sentence. The failure has a dollar figure attached to it. The moment a scenario names a financial consequence, a gate has already won before you finish reading the four options. Second, domain three. A codebase has react components, a p i handlers, and database models, each with their own conventions. Test files sit scattered next to the code they test, rather than gathered in one place. The guide asks how to make sure Claude Code applies the right convention automatically regardless of location. The published answer is a rules directory with YAML frontmatter glob patterns, matching files by name rather than by folder. A root CLAUDE.md relying on inference is offered, and so is a directory level file that cannot reach scattered files. Both fall short for the reason episode six spent fifteen minutes on. A glob does not care where a file lives. Only what it is. That is the whole difference between a directory-shaped solution and a pattern-shaped one, and it is worth carrying past this exam entirely. Third, domain four. A team wants to cut API costs. Move two workflows onto the Message Batches API for its cost savings. A blocking check that has to finish before a developer can merge. An overnight report nobody reads until morning. The guide asks how to evaluate that proposal. The published answer batches only the overnight report. The blocking check stays on real time calls. Batch carries no guaranteed latency, and no guaranteed latency is precisely the property a blocking workflow cannot tolerate. Moving both to batch, or adding a timeout fallback, both miss that the discount is not the variable that matters here. Pattern three’s other half: know what you are actually optimizing for before you reach for the cheaper option. A discount that costs you a blocked engineer is not a discount. It is a cost with a different name on it. Three questions, three domains, the same five patterns underneath every one of them. That is not a coincidence. It is the blueprint working exactly as designed. It is also the honest reason a foundations exam rewards judgment more than memorization. You cannot memorize your way to a pattern. You can only recognize one, and recognizing one is exactly what eleven episodes of worked questions were for. One last practical note, and then we are done. If it goes well, you have twelve months. Renewal is a free, non-proctored review, rather than sitting the whole thing again, as long as you renew on time. If it does not go well, the wait is fourteen days before a second attempt. Thirty after a second failure. Ninety after a third. The fee is due again each time. Fail once, and the clock resets in two weeks, not two years. Compare that against most professional credentials. A miss there costs a full year, and a second application fee, before you even get another attempt. This exam is built to let you try again, not to punish a single bad morning permanently. Neither outcome is the end of anything. A scaled score is a measurement, not a verdict. The guide itself says as much. Pass or fail against a fixed standard. Never against how anyone else in the room happened to do that day. That’s the series. Thirteen episodes, five domains, thirty task statements. Built entirely from a public guide. And from the kind of hands on work anyone running Claude Code every day already does, without calling it studying. If it helped, the disclaimer at the front of every episode is still true at the back of the last one. That repetition, thirteen times over, was never filler. It was the promise the whole series was built to keep. Independent, unofficial. Not affiliated with or endorsed by Anthropic. A companion to the real guide, never a replacement for reading it yourself. Good luck in the room. This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit brodynetworks.substack.com (https://brodynetworks.substack.com?utm_medium=podcast&utm_campaign=CTA_1)

4 Sept 2026
Episode 11. Reliability, Escalation, Errors, Review, Provenance
Domain 5, task statements 5.2, 5.3, 5.5 and 5.6. Fifty five percent resolution against an eighty percent target, because an agent escalated the easy cases and attempted the hard ones. The three real escalation triggers, structured error context across multi-agent systems, why an aggregate accuracy number can hide a real problem, and keeping a claim attached to its source through synthesis. Independent and unofficial. This series is not affiliated with, sponsored by, or endorsed by Anthropic. Nothing in it is exam content. Every practice question was written for this show against the published exam guide, which is a free public document linked below. Chapters * 0:00 Cold open and disclaimer * 0:22 Fifty five percent against an eighty percent target * 1:01 The three legitimate escalation triggers * 1:49 Why sentiment and confidence fail as proxies * 2:31 Honoring an explicit request immediately * 3:07 Escalating on a genuine policy gap * 3:28 Multiple matches: ask, do not guess * 3:58 Structured error context across agents * 4:06 Access failure versus a valid empty result * 4:53 The two anti-patterns * 5:40 Why an aggregate accuracy number can hide a problem * 6:00 Stratified sampling of high-confidence extractions * 6:45 Field-level confidence calibration * 6:50 Provenance and claim-source mappings * 7:38 Annotating conflicting statistics * 7:47 Dates against false contradictions * 8:47 Worked question * 9:35 Why the other three options are there * 10:10 What the exam will ask, and next time Sources * Claude Certified Architect, Foundations Exam Guide, Version 1.0 (section 6, task statements 5.2, 5.3, 5.5 and 5.6) (https://everpath-course-content.s3-accelerate.amazonaws.com/instructor%2F6nizmqk8tpzpfjvt6qmmav7rh%2Fpublic%2F1783542750%2FClaude+Certified+Architect+%E2%80%93+Foundations+Exam+Guide.pdf) Transcript This is Passing C C A R F, an independent study companion for the Claude Certified Architect, Foundations exam. Sixty items, a hundred and twenty minutes, and a published blueprint that tells you almost exactly what it is going to ask. Independent and unofficial. Not affiliated with or endorsed by Anthropic. No exam content. Fifty five percent resolution against an eighty percent target. Not because the agent was attempting cases it could not handle. It was escalating the easy ones and attempting the hard ones. Almost exactly backwards. Nobody had told it what escalation was actually for. Domain five continues. Task statements five point two, five point three, five point five and five point six. Escalation, error propagation across agents, honest measurement, and keeping a claim attached to its source. The last stretch of the exam blueprint, and one of the highest density episodes in the series. Start with escalation, because most systems get the trigger wrong before they get anything else wrong. There are exactly three legitimate reasons to escalate. The customer explicitly asks for a human. Policy is genuinely silent or ambiguous on the specific request in front of you. The agent cannot make meaningful progress, not that it does not want to, that it actually cannot. Notice what is missing from that list. Complexity is not on it. A case being hard is not, by itself, a reason to hand it to a human. The system I opened with had inverted this exactly. It was escalating on some proxy for difficulty, and treating everything else as safe to attempt. Difficulty was never the right signal in the first place. Sentiment and self reported confidence are the two proxies people reach for anyway. The exam wants you to know precisely why both fail. An upset customer with a completely straightforward request does not need a human. Confidence describes how sure the model feels, and a model can feel very sure while being completely wrong. Neither number measures the thing you actually care about, which is whether this specific case matches one of the three real triggers. A calm customer with an unanswerable request needs a human just as much as a furious one does. Tone tells you almost nothing about which of the three triggers actually applies. There is a nuance inside the first trigger worth holding onto. A customer who explicitly demands a human gets one immediately, no investigation first, full stop. A customer who is upset but has not demanded escalation gets acknowledgment and an offer to resolve. It escalates only if they say so again. Read frustration as an automatic escalation trigger, and you hand off cases the agent could have closed in one turn. Ignore an explicit request for a human, and you have ignored the one trigger that needs no further judgment at all. The policy gap trigger deserves its own moment, because it is not the same as a hard question. Competitor price matching, when your policy only ever addresses your own site’s price adjustments, is not a difficult case. It is a case with no answer written down anywhere. The correct move is escalating precisely because policy is silent. Not attempting a guess and hoping it lines up with whatever the company would have wanted. One more pattern worth naming: multiple matches on a lookup. Two customers with the same name, three orders matching a partial description. The fix is asking for another identifier, not picking the most likely match and hoping. A heuristic guess here is a guess about someone’s account. Getting it wrong is not a near miss. It is the wrong customer’s data in the wrong conversation. One extra question costs thirty seconds. A wrong match costs someone else’s private information landing in a stranger’s conversation. Now task statement five point three. This is domain two’s error work, one level up. Across a multi agent system, instead of inside one tool. Structured error context, failure type, what was actually attempted, partial results, alternatives. That is what lets a coordinator make an actual recovery decision, instead of guessing blind. The same access-failure-versus-empty-result line from a few episodes back matters even more here. A coordinator sitting above several subagents cannot afford to confuse the two, even once. A subagent reporting a timeout needs a different coordinator response than a subagent reporting a genuinely empty, successful search. Collapse those into one undifferentiated failure signal, and the coordinator loses the information it needs to respond correctly to either one. A retry makes sense for one. A retry is wasted effort for the other. From the coordinator’s seat, undifferentiated failure looks identical either way, until something tells it otherwise. Generic error statuses hide exactly this distinction. Search unavailable tells the coordinator nothing about what to do next. Two anti patterns bracket the correct behavior on either side. Silently swallowing an error and reporting empty results as if the query simply found nothing is one. Killing an entire multi agent workflow because one subagent failed is the other. The right answer sits between them. Local recovery inside the subagent for anything transient. Structured context upward for anything that could not be resolved locally. The overall workflow continues on partial results, with the gap honestly annotated, rather than hidden or treated as fatal. Task statement five point five is where honest measurement lives. It opens with a number that sounds reassuring, and is not telling you what you think. Ninety seven percent overall accuracy can hide a document type performing at sixty percent. Buried inside an average that never breaks the number down by segment. Stratified random sampling is the fix. It targets exactly the extractions a naive process would never think to check. The high confidence ones. Confidence is the model’s own self assessment, and self assessment is not ground truth. Sampling across confidence levels, not just auditing the cases the model already flagged as uncertain. That is what actually catches an error pattern the model does not know it is making. Analyzing accuracy by document type and by field, before you reduce human review on the strength of one aggregate number. This is not optional caution. It is the only way to know whether that ninety seven percent is real everywhere. Or real on average and false somewhere specific. An aggregate number is an average of many truths and a few failures. The average alone cannot tell you which document type is hiding the failures. Calibrate field level confidence against labeled validation sets, and route by that calibration. Low confidence and contradictory source documents both belong in front of a human, prioritized by the reviewer time you actually have. Reviewer hours are finite. Spend them where the model is actually unsure, not spread evenly across a pile that mostly did not need a second look. Last task statement: five point six, provenance. When multiple sources feed one synthesis, source attribution is exactly what summarization steps lose first. The same way task statement five point one’s numbers get lost first in a single conversation. A claim survives. Which source said it does not. Not unless something forces it to travel alongside the claim structurally, rather than trusting prose to carry it. The fix is a structured claim-source mapping the synthesis agent is required to preserve, not paraphrase, all the way through. When two credible sources disagree on a statistic, the answer is never picking one and moving on. Annotate the conflict. Attribute both values to their sources. Let whoever reads the report see the disagreement, rather than a single number that quietly erased it. Dates matter here in a specific way. Two sources reporting different numbers six months apart are not necessarily contradicting each other. They may simply be describing different moments. Requiring a publication or collection date on every structured finding is what lets a reader, or a coordinator, tell a real contradiction apart. From an honest change over time. One closing point on how a synthesis actually presents what it found: match the shape of the content to the content itself. Financial data as a table. News as prose. Technical findings as a structured list. Forcing everything into one uniform format for the sake of consistency loses the structure. The structure that made the original data legible in the first place. A table of numbers forced into a paragraph is harder to scan, not easier. Consistency is not the same virtue as clarity, and the two occasionally pull in opposite directions. Let me work a question. A synthesis report combines findings from four sources on a company’s revenue growth. Two sources report a nineteen percent increase. One reports twelve percent. One reports twenty two percent. The published report shows a single number, nineteen percent, with no mention of the other two figures. What is wrong, and what is the fix. Option one, nothing is wrong, since nineteen percent is what most sources agree on. Option two, the synthesis should have annotated the conflict, presenting all three values with source attribution rather than selecting the majority figure. Option three, the report should have averaged all four numbers into one. Option four, the report should have used only the most recent source and discarded the rest. The answer is option two. Three different reported figures on the same metric is a genuine conflict. Arbitrarily selecting one, even a majority one, erases information a reader might need. Especially if the discrepancy traces to different measurement periods or methodologies. Option one treats agreement by count as agreement on truth. Two sources repeating the same number could just as easily share a common origin, rather than independent confirmation. Option three manufactures a number nobody actually reported. That is worse than picking a real one. An average of four different measurements is not a measurement of anything. Option four assumes recency resolves the conflict. Nothing in the question establishes that the sources even carry different dates. Discarding real data on an unverified assumption is exactly the arbitrary selection this task statement warns against. Three sentences on what the exam will ask. Expect an escalation scenario where the correct trigger is one of exactly three, and complexity alone is always the wrong answer. Expect a structured error context question. The coordinator’s correct decision depends on telling an access failure apart from a valid empty result. And when synthesis combines conflicting sources, the answer is almost always annotate and attribute both, never average, never pick a favorite. Next time, the final episode. Exam day, the room, the rules, and the five patterns that decide most of the questions you will actually see. This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit brodynetworks.substack.com (https://brodynetworks.substack.com?utm_medium=podcast&utm_campaign=CTA_1)

4 Sept 2026
Episode 10. Context That Survives
Domain 5, task statements 5.1 and 5.4. An agent quietly loses a customer’s exact refund amount to progressive summarization. What summarization eats first, the lost-in-the-middle effect, trimming tool output at the source, and keeping a long codebase exploration from degrading into generic answers. Independent and unofficial. This series is not affiliated with, sponsored by, or endorsed by Anthropic. Nothing in it is exam content. Every practice question was written for this show against the published exam guide, which is a free public document linked below. Chapters * 0:00 Cold open and disclaimer * 0:21 The refund amount that got summarized away * 0:59 What summarization eats first * 1:52 The lost-in-the-middle effect * 2:30 Tool output outpacing relevance * 3:14 The case facts block * 3:51 Ordering aggregated inputs * 4:22 Trimming at the source * 4:52 Passing complete conversation history * 5:26 Structured facts over reasoning chains * 6:03 Context degradation in long exploration * 6:56 Scratchpad files * 7:17 Delegating discovery to a subagent * 7:50 Summarizing between exploration phases * 8:20 Structured state manifests for crash recovery * 8:54 /compact * 9:08 Worked question * 10:20 Why the other three options are there * 11:12 What the exam will ask, and next time Sources * Claude Certified Architect, Foundations Exam Guide, Version 1.0 (section 6, task statements 5.1 and 5.4) (https://everpath-course-content.s3-accelerate.amazonaws.com/instructor%2F6nizmqk8tpzpfjvt6qmmav7rh%2Fpublic%2F1783542750%2FClaude+Certified+Architect+%E2%80%93+Foundations+Exam+Guide.pdf) * Context windows (https://docs.claude.com/en/docs/build-with-claude/context-windows) Transcript This is Passing C C A R F, an independent study companion for the Claude Certified Architect, Foundations exam. Sixty items, a hundred and twenty minutes, and a published blueprint that tells you almost exactly what it is going to ask. Independent and unofficial. Not affiliated with or endorsed by Anthropic. No exam content. The agent forgot the refund amount from turn three. Not the whole conversation. Just the one number that actually mattered, the three hundred and forty dollars the customer had already been promised. By turn nine, it was offering a different figure, confidently. Somewhere along the way, that first number got summarized down to about three hundred dollars. About is not a number you can issue a refund against. Nobody rewrote that number on purpose. It just got compressed, one summary at a time. Until the version left standing was close enough to sound right, and wrong enough to matter. Smaller than the others, and no easier to lose points on, because reliability failures are the ones customers actually notice. Domain five, context management and reliability, fifteen percent of the exam. Task statements five point one and five point four today. What summarization actually eats first, and how to keep a long exploration from quietly losing the plot. Start with what gets lost, because it is not random. Progressive summarization is good at compressing prose and terrible at protecting numbers. Amounts, percentages, dates, a specific thing the customer said they expected. Those get rounded, softened, and generalized every time a summary compresses them. Three hundred and forty quietly becomes about three hundred. That is not a rounding error. It is a different number. There is a second failure sitting right next to it, and it is about position, not compression. Long inputs get read unevenly. The beginning gets attention. The end gets attention. The middle is where things go missing, not because the model cannot see it, but because it reliably weights it less. A key finding buried in the middle of a long aggregated report is more likely to get skipped. Statistically, compared to the same finding sitting first or last. Nothing about the finding changed. Only where you put it did. That is the entire fix, and it costs nothing more than deciding where a sentence goes. Third failure, and this one is just arithmetic: tool output piles up faster than its relevance does. An order lookup can return forty or more fields. Maybe five of them matter for the question actually in front of you. Every one of those unused fields still sits in context. It still costs tokens. It still competes for attention against the five that actually count. Thirty five fields nobody asked about are not free just because nobody reads them. They are still there, still competing, on every single turn. Multiply that by every order lookup in a long session. The waste compounds far faster than anyone tracking it turn by turn would expect. The fix for the first failure has a name worth knowing: a case facts block. Pull the transactional numbers, amounts, dates, order numbers, statuses, out of the narrative entirely. Hold them in a structured block, included in every prompt, sitting outside whatever gets summarized. Summarization can compress the story. It should never get anywhere near the three hundred and forty dollars. Let the narrative get shorter every turn if it has to. The number does not get to shrink with it. Everything else in that conversation is allowed to drift a little in the retelling. That one figure is not. The fix for the middle problem is about order, not content. Put key findings first in an aggregated input, not buried wherever they happened to land. Use explicit section headers so the model has structural cues, not just prose, guiding it to what matters. You are not fighting the lost in the middle effect by writing more clearly. You are fighting it by not putting anything you care about in the middle to begin with. Move the finding, not the sentence around it. The fix for field bloat is trimming at the source. Keep the five relevant fields from that order lookup. Drop the other thirty five before they ever accumulate in context, not after. Trimming after the fact means you already paid the token cost and already diluted attention with everything you were about to discard. Trim at the boundary, the moment the tool result comes back. The thirty five unused fields never get the chance to cost you anything at all. There is a basic requirement underneath all of this, easy to assume rather than actually verify. The complete conversation history has to go back in every subsequent request. The API does not remember anything between calls on its own. Drop part of that history by accident, trying to save tokens, and you have not trimmed context. You have broken the model’s ability to reason coherently about anything that came before the gap. Trimming and forgetting look identical from the outside, right up until the model needs the thing you dropped. One more piece belongs with task statement five point one, and it connects straight back to two episodes ago. Say an upstream agent hands off to a downstream one with a tight context budget. Send structured facts and citations. Not verbose reasoning chains. A downstream agent with limited room does not need to see how you arrived at a conclusion. It needs the conclusion, with enough metadata, dates, sources, to actually use it correctly. Send the reasoning anyway, and you have not helped the downstream agent. You have just spent its limited budget on your own thinking, instead of its own. Now task statement five point four. It is the same problem, at the scale of a whole investigation instead of one conversation. Long codebase exploration degrades in a specific, recognizable way. The agent starts giving answers about typical patterns instead of the actual class it found forty minutes ago. That is not the model getting worse. It is the model losing track of specifics it already established, and quietly substituting a generic default instead. A default sounds plausible. It sounds like it belongs. It is also not the actual class. Confidently generic is a worse failure than an honest I don’t remember would have been. At least an honest gap tells you to go look again. A confident, wrong default tells you nothing is wrong at all. Scratchpad files are the direct countermeasure. Have the agent write down key findings as it discovers them. Do not just hold them in a fading context window. Reference that file for later questions, instead of relying on memory of a conversation that keeps growing. A written fact does not degrade the way a summarized one does. Delegation helps here too, the same Explore subagent idea from the CI episode, applied to investigation instead of implementation. Spawn a subagent to answer one specific question. Find every test file. Trace the refund flow’s dependencies. The main agent holds only the high level picture. The verbose legwork happens somewhere it cannot dilute the coordination layer. The main agent never sees every file the subagent opened along the way. It sees the answer to the one question it actually asked. Between phases, summarize what the first phase actually found before spawning subagents for the next one. Inject that summary into their starting context. Each phase should start knowing what the last one learned, not starting cold and rediscovering it. Rediscovery is not thoroughness. It is the same work, paid for twice. A summary handed forward costs a few sentences. Rediscovering it later costs the whole exploration again. For anything long running enough to actually crash, structured state exports matter. Each agent writes its state to a known location. The coordinator, on resume, loads a manifest and injects it back into the relevant prompts. A crash then costs you the time since the last checkpoint, not the whole investigation from scratch. Six hours of exploration surviving a crash intact is the entire point of writing the manifest in the first place. Skip that step, and a crash at hour five is not a delay. It is the whole day, starting over. And when context is genuinely filling with exploration you no longer need, slash compact reduces it directly. It does not let an already strained context window get worse, turn after turn. Let me work a question. A coordinator running a six hour codebase migration keeps its full exploration history in context the entire time. Around hour four, it starts describing files it examined earlier using generic language. Typical service pattern. Standard handler. Instead of the specific details it found and reported correctly at hour one. What is the most direct fix. Option one, increase the context window size so more history fits without being summarized. Option two, have agents write key findings to scratchpad files as they go. Reference those files instead of relying on the accumulating conversation. Option three, restart the migration from scratch with a cleaner prompt. Option four, instruct the agent to be more precise and specific in its language. The answer is option two. The specifics were established correctly at hour one. They did not survive being carried passively in a growing context. A written record outside that context is what actually preserves them. A conversation can drift. A file on disk does not. Option one treats this as a capacity problem. A bigger window still summarizes eventually. It just delays the same degradation by some number of hours, rather than fixing what causes it. Option three throws away four hours of legitimate progress to solve a problem that scratchpad files solve without losing any of it. Option four asks for more precision from a model that already knew the specific details once. The information did not vanish because the wording got sloppy. It got generalized because nothing preserved it outside the conversation that was quietly compressing around it. Ask a model to sound more precise about a fact it no longer clearly holds, and it will sound precise anyway. That is the trap. Three sentences on what the exam will ask. Expect a scenario where a specific number degrades into a vague one over a long conversation. The fix is a persistent facts block outside summarized history. Never a better summarization prompt. Expect a codebase exploration question where the tell is generic language replacing specific findings from earlier. The answer is scratchpad persistence, not a bigger context window. And when crash recovery comes up, know the shape: each agent exports state, the coordinator loads a manifest on resume. Next time, the second half of domain five. Fifty five percent resolution against an eighty percent target. And why an agent escalating the easy cases and attempting the hard ones is not a training problem. This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit brodynetworks.substack.com (https://brodynetworks.substack.com?utm_medium=podcast&utm_campaign=CTA_1)

4 Sept 2026
Episode 9. Structured Output, Schemas, Validation, Batch
Domain 4, task statements 4.3, 4.4 and 4.5. An extraction pipeline invents a clean, confident, entirely fake invoice number. Why a valid schema does not mean a true answer, nullable fields and enum escape hatches as fabrication controls, when a retry can and cannot fix a bad extraction, and when the Message Batches API saves money instead of costing you a blocked engineer. Independent and unofficial. This series is not affiliated with, sponsored by, or endorsed by Anthropic. Nothing in it is exam content. Every practice question was written for this show against the published exam guide, which is a free public document linked below. Chapters * 0:00 Cold open and disclaimer * 0:23 The invented invoice number * 0:59 Tool use with JSON schemas * 1:33 Syntax errors versus semantic errors * 2:23 Nullable fields against fabrication * 3:01 Enum escape hatches: unclear and other * 3:39 tool_choice any versus forced, for extraction * 4:15 Retry with the specific validation error * 4:40 Where retry stops working * 5:20 Self-checking schemas: totals and conflicts * 6:08 The detected-pattern field * 6:29 The Message Batches API * 6:54 Matching batch to a workload’s patience * 7:33 The multi-turn tool calling limit * 7:58 Sizing submission windows to an SLA * 8:30 custom_id and resubmitting failures * 8:52 Testing prompts on a sample first * 9:16 Worked question * 10:09 Why the other three options are there * 11:10 What the exam will ask, and next time Sources * Claude Certified Architect, Foundations Exam Guide, Version 1.0 (section 6, task statements 4.3, 4.4 and 4.5) (https://everpath-course-content.s3-accelerate.amazonaws.com/instructor%2F6nizmqk8tpzpfjvt6qmmav7rh%2Fpublic%2F1783542750%2FClaude+Certified+Architect+%E2%80%93+Foundations+Exam+Guide.pdf) * Structured outputs (https://docs.claude.com/en/docs/build-with-claude/structured-outputs) * Message Batches API (https://docs.claude.com/en/docs/build-with-claude/batch-processing) Transcript This is Passing C C A R F, an independent study companion for the Claude Certified Architect, Foundations exam. Sixty items, a hundred and twenty minutes, and a published blueprint that tells you almost exactly what it is going to ask. Independent and unofficial. Not affiliated with or endorsed by Anthropic. No exam content. The extractor pulled an invoice number off a document that never had one. Not garbled. Not close. A clean, confident, entirely invented number, sitting in the output field exactly where a real one belonged. The schema was valid. The JSON parsed perfectly. The number was fiction. Domain four continues. Task statements four point three, four point four and four point five today. How to actually guarantee structured output. What to do when it comes back wrong. When batch processing is the right tool, and when it is a trap. Start with the fix for syntax, because it is a real fix, and it is not the whole fix. Tool use with a JSON schema is the reliable route to structured output. Define the schema as a tool’s input, and the model returns data that already matches it. No more hand parsing a wall of prose, hoping the brackets close where you expect them to. That old approach failed in a very specific way. The model got the content right and the format wrong. A regex somewhere downstream choked on a comma it did not expect. Here is the boundary you need to hold in your head, because the exam draws it sharply. Strict schemas eliminate syntax errors. They do not eliminate semantic errors. A schema happily accepts a total that does not match the sum of its own line items. It accepts a value sitting in the wrong field entirely, formatted correctly, meaning nothing. Syntax and truth are different problems, and a schema only ever solves the first one. A well formed lie is still a lie. The schema just makes sure it is a well formed one. That gap is exactly where the invented invoice number lives. Nothing about a well formed schema stops a model from confidently filling a required field with something plausible. Not when the real value simply is not in the source document. The direct countermeasure: make fields optional, not required, whenever the source document might not contain the information. A nullable field can come back empty. A required field cannot. A model asked for something that must exist, and does not, will sometimes invent rather than admit absence. Optionality is not a weaker schema. It is the schema being honest about what a real document can and cannot promise. A required field is a promise the source document did not actually make. Do not force the model to keep a promise you made on the document’s behalf. Enums close a related gap. Real documents have messy middles, categories that almost fit. Forcing a fixed enum with no escape hatch pushes the model to pick the closest wrong answer, rather than say so. Add unclear as a legitimate value. Add an other option paired with a free text detail field. Now ambiguity has somewhere honest to go, instead of getting silently rounded off to whichever category happened to be nearby. A category with no honest exit does not remove ambiguity. It just hides it inside a wrong answer that looks exactly like a right one. Tool choice comes back into play here, and this time in the extraction context specifically. Set it to any when you are working across multiple possible extraction schemas. Use it when you do not yet know which one this document actually needs. That guarantees a tool gets called, without locking in which one ahead of time. Force a specific tool by name when a step genuinely has to run first. Extract metadata forced ahead of any enrichment step. Same ordering logic from a few episodes back, just applied here to extraction instead of research. Now task statement four point four, and this is where validation actually earns its keep or reveals it never had any. When a structured output fails validation, the fix is not a blind retry. It is a retry that includes the original document, the failed attempt, and the specific validation error. The model gets something concrete to correct, rather than a second blind guess. But know exactly where that retry stops working, because this is the single most testable idea in the task statement. A format error can be fixed on retry. A structural mismatch can be fixed on retry. Information that was never in the source document cannot be fixed on retry. No matter how many times you ask. There is nothing there to find. Recognizing which failure you are looking at, before you spend a second call on it, is the actual skill. Send a format error back for retry, and the model usually fixes it. Send an absence back for retry, and you have paid for a second identical guess. There is a self correcting schema pattern worth knowing by name. It catches the invented number problem at the design level, instead of after the fact. Extract a calculated total alongside a stated total, and flag it when they disagree. Add a conflict detected boolean for exactly this kind of internal contradiction. Two numbers that should match and do not is not noise. It is the document telling you something is wrong before a human ever has to notice by hand. The schema is not just capturing data anymore. It is capturing whether the document agrees with itself. That is precisely the signal that would have caught a fabricated invoice number. The moment it failed to reconcile with the line items sitting right next to it. One more habit worth building in from day one: a detected pattern field on every finding. When developers dismiss a flagged issue, that field lets you analyze which patterns keep generating false positives. Systematically. Instead of a vague sense that the tool is noisy, without ever being able to say exactly where. Last task statement: the Message Batches API. The trap the exam is watching for is not technical. It is a manager’s instinct. Batch cuts cost roughly in half. It runs within a processing window up to twenty four hours. It carries no guaranteed latency SLA, which is the detail that decides everything else in this task statement. Match the tool to the workload’s actual patience, not to the discount. An overnight report has all night. A weekly audit has all week. Nightly test generation has until morning. These tolerate the twenty four hour window without anyone noticing it happened. A pre-merge check has an engineer sitting there, blocked, waiting on the result before they can ship. Batch it, and you have traded a small cost saving. For an engineer staring at a spinner with no idea when it will resolve. Half price on a merge that never happens is not a discount. It is a stalled team. One mechanical limit worth knowing cold: batch does not support multi turn tool calling inside a single request. No mid-request tool execution and return. Say your workflow genuinely needs that back and forth. Batch is not an option for it, regardless of how patient the workload otherwise is. No amount of patience buys back a capability the API does not offer inside a batch request. Sizing submission windows against a real SLA is arithmetic, not guesswork. Promise a thirty hour turnaround. A twenty four hour processing window means you can only submit every four hours or so. That still lands inside your own promise, with room to spare. Do the math against the SLA you actually made, not the SLA you hope batch delivers on a good day. Batch has no obligation to move faster just because you are counting on it to. Custom ID fields are what make failure handling tractable at scale. Say some documents in a batch of a hundred fail. Resubmit only the failures, identified by their custom ID, with whatever fix the failure actually needs. Chunk an oversized document. Do not resubmit the entire hundred blind. And before any of that: refine your prompt against a small sample first. Test on ten documents before you commit to submitting ten thousand. A prompt that is seventy percent right at sample size is a warning, not a surprise you want at scale. Discover that after paying for the full batch, and you are resubmitting three thousand failures, one expensive cycle at a time. Let me work a question. A team extracts structured data from scanned invoices. On invoices missing a purchase order number, roughly one in six, the extractor returns a plausible looking number. Instead of leaving the field empty. Retrying those extractions with the same prompt produces the same fabricated numbers every time. What is the correct fix. Option one, keep retrying with more explicit instructions to only extract what is actually present. Option two, make the purchase order field nullable instead of required. The model can then return null when the information is genuinely absent. Option three, increase the model’s temperature to encourage more varied, hopefully more honest responses. Option four, add a longer few shot section showing correct purchase order extraction. The answer is option two. A required field asked for information that does not exist is exactly the condition that produces fabrication. Retrying does not help. The underlying document has not changed, and never will. Option one is a retry with extra words attached. This task statement is explicit that retries do not fix an absence of information. No matter how the instruction is phrased around it. Option three treats fabrication as a randomness problem. It is a schema problem, a field structurally unable to admit that the answer is missing. More variance does not add honesty, it adds noise on top of the same underlying design flaw. Option four teaches the model to extract when a purchase order number is genuinely present. It teaches nothing about correctly reporting absence, which is the actual failure here. More examples of the easy case do not touch the hard one. Three sentences on what the exam will ask. Expect a scenario where structured output is syntactically perfect and semantically wrong. The fix is schema design: nullable fields, enum escape hatches, self checking totals. Never a retry. Expect a retry-versus-fix question where the tell is whether the missing information could ever have existed in the source. That alone decides whether retrying helps, or wastes a call. And expect a batch-versus-synchronous question that turns entirely on whether the workflow can tolerate an unpredictable wait, never on the discount alone. Next time, domain five, and the first of the reliability episodes. An agent that forgets a refund amount from three turns ago, and why summarizing harder only makes it worse. This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit brodynetworks.substack.com (https://brodynetworks.substack.com?utm_medium=podcast&utm_campaign=CTA_1)

4 Sept 2026
Episode 8. Precision Prompting, Criteria, Few Shot, Multi Pass
Domain 4, task statements 4.1, 4.2 and 4.6. A review bot with a high false positive rate gets turned off by the team that built it. Why be conservative fails as an instruction, what makes a few shot example actually teach judgment instead of pattern matching, and why the same session that wrote code is the wrong session to review it. Independent and unofficial. This series is not affiliated with, sponsored by, or endorsed by Anthropic. Nothing in it is exam content. Every practice question was written for this show against the published exam guide, which is a free public document linked below. Chapters * 0:00 Cold open and disclaimer * 0:23 The review bot the team turned off * 1:01 Why be conservative fails * 1:44 Explicit categorical criteria * 2:38 Named categories over confidence thresholds * 3:51 Severity levels with code examples * 4:32 Few shot: examples are for ambiguity * 4:53 Reasoning versus answer keys * 5:29 Format consistency via examples * 5:50 Tool selection and messy documents * 6:52 Why self review is structurally biased * 7:40 The fix: an independent instance * 8:05 Multi pass: per-file and cross-file * 8:50 Confidence alongside findings * 9:16 Worked question * 10:22 Why the other three options are there * 11:09 What the exam will ask, and next time Sources * Claude Certified Architect, Foundations Exam Guide, Version 1.0 (section 6, task statements 4.1, 4.2 and 4.6) (https://everpath-course-content.s3-accelerate.amazonaws.com/instructor%2F6nizmqk8tpzpfjvt6qmmav7rh%2Fpublic%2F1783542750%2FClaude+Certified+Architect+%E2%80%93+Foundations+Exam+Guide.pdf) * Few shot prompting (https://docs.claude.com/en/docs/build-with-claude/prompt-engineering/multishot-prompting) Transcript This is Passing C C A R F, an independent study companion for the Claude Certified Architect, Foundations exam. Sixty items, a hundred and twenty minutes, and a published blueprint that tells you almost exactly what it is going to ask. Independent and unofficial. Not affiliated with or endorsed by Anthropic. No exam content. The team turned the review bot off. Not because it never caught anything real. Because for every real bug, it flagged four things that were not bugs at all. Engineers stopped reading its comments. A tool that is wrong most of the time trains people to ignore it, including the times it happens to be right. Domain four, prompt engineering and structured output, twenty percent of the exam, tied with domain three. Task statements four point one, four point two and four point six today. Precision, few shot examples, and how to structure a review so it does not fall apart at scale. Start with the fix somebody tried first, because it is the fix almost everyone tries first, and it does almost nothing. Be conservative. Only report high confidence findings. Those instructions sound like precision. They are not criteria. They are vibes with a confident tone. The model has no fixed definition of conservative to apply consistently. It applies its own shifting sense of it, turn to turn. Ask ten reviewers what conservative means and you get ten thresholds. Ask the model the same vague question ten times, across ten different files, and you get the same drift. What actually works is explicit, categorical criteria. Not check that comments are accurate. Flag a comment only when the claimed behavior contradicts what the code actually does. That is a test the model can run the same way every time. It names a specific condition, rather than a feeling. One names a condition. The other names a mood. Here is why this matters more than it sounds like it should. False positives do not stay contained to the category that produced them. A high false positive rate in one category burns trust in every category, including the ones the tool gets right constantly. Once someone stops reading the output, accuracy in the categories they never see does not help them at all. A correct finding nobody reads has the same practical value as no finding at all. Two moves fix this in practice. First, write specific criteria that state what to report and what to explicitly skip. Bugs and security issues, report. Minor style preferences, local team patterns, skip. Not confidence thresholds. Actual categories, named. A confidence threshold asks the model to rate its own certainty. Certainty is exactly the thing an overconfident model gets wrong most often. A named category asks it to check a fact instead. Second, and this is the one people resist, because it feels like giving something up. If one category is genuinely noisy, turn it off. Temporarily disable the worst offender while you fix its prompt, rather than let it keep eroding trust in the categories working fine. A team will forgive a tool for not covering everything. They will not forgive a tool that cries wolf on every single review. A gap is honest. Everyone knows the tool has not looked there yet. Coverage lost is a gap you can fill later. Trust lost is a much longer road back, and most teams never bother walking it. They just stop reading the bot. Severity classification gets the same treatment. Do not just say rate this critical, high, or low. Give each severity level a concrete code example of what belongs there. A shared example is worth more than a shared adjective. Two people, or two runs of the same model, read an adjective differently. They read an example the same way. Show what critical looks like in actual code, a hardcoded credential, an unbounded loop that can hang a request. Show what low looks like the same way, a missing comment, a variable name that could be clearer. The adjective invites debate. The example ends it. Now few shot prompting, task statement four point two, and here is the core idea worth holding onto: examples are for ambiguity. Say detailed instructions alone still produce inconsistent output. That is the signal you need two to four targeted examples, not a longer paragraph of more instructions. The examples that actually help are not just answer keys. They show reasoning. An example that demonstrates why one tool was chosen over a plausible alternative teaches the model something durable. It generalizes that judgment to a case it has never seen. An example that just shows an answer with no reasoning attached teaches pattern matching to that one case, and nothing else. Bare answers train a lookup table. Reasoning trains a judgment. One scales to new cases. The other only ever repeats the old ones. Format consistency is the other place few shot earns its keep. Show two or three examples in the exact output shape you want. Location, issue, severity, suggested fix. The model locks onto that shape far more reliably than a written description of the same fields ever manages. Few shot also earns its place wherever a request is genuinely ambiguous rather than just varied. Tool selection between two plausible tools is one such case. Show two or three examples of a borderline request. Pair each with the tool that was actually chosen, and a line of reasoning for why. The model is not memorizing those exact requests. It is learning the shape of the decision. That is what lets it handle the next borderline case correctly, the one none of your examples ever showed it. There is a sharper use case worth knowing by name: extraction from messy documents. Vary the examples across the kinds of structural variety you actually expect. Inline citations against full bibliographies. A methodology section against details buried in a table. Do that, and empty or null extractions on the fields that matter drop hard. The model has seen the shape of the mess before, not just the shape of a clean case. Last task statement, and it belongs here because it is really about trust in a different form: the review itself. Task statement four point six, multi instance and multi pass architectures. Here is the limitation the exam wants you to know cold. A model that just generated some code keeps the reasoning it used to write that code, in the same session. Ask it to review its own work, and it is reviewing decisions it already committed to. With the confidence of having just made them. Telling it to think harder, or extending its reasoning time, does not remove that bias. It is still the same session, holding the same committed context. Nothing about thinking longer changes whose decisions are under review. The fix is not a smarter self review prompt. It is a second, independent instance, one with no memory of writing the code, looking at it fresh. That independent instance catches things the generator does not. Not because it is smarter. Because it never had a reason to defend the first draft’s decisions. It walks in cold. Nothing to defend, nothing already decided. Multi pass review solves a different problem, the one from episode one, back again in this domain. Split a large review into a per file pass for local issues. Add a separate cross file pass for how the pieces interact. Do both in one sweep across everything, and attention dilutes exactly the way it did with the fourteen file pull request. Split it, and each pass actually gets to focus. One pass reads a single file closely. A second pass reads across files for exactly the kind of contradiction a narrow, file by file view would never surface. Neither pass is doing the other’s job. That is exactly why both together outperform one pass trying to do everything at once. One more piece worth naming: confidence alongside findings. Have the review self report a confidence level on each individual finding, not just an overall verdict. That number is what lets you route intelligently later. Low confidence findings go to a human. High confidence ones go straight through. Nobody treats every finding the same, regardless of how sure the model actually was. Let me work a question. A code review tool flags issues with instructions telling it to be thorough and to use good judgment about what matters. Two different files with nearly identical patterns get inconsistent verdicts, flagged in one, ignored in the other. What is the most direct fix. Option one, tell the model to be more consistent across files. Option two, replace the judgment based instruction with explicit criteria. Name which patterns to flag and which to skip, plus two or three examples showing the reasoning. Option three, run the review three times and take the majority verdict. Option four, increase the model’s context window so it can see both files at once. The answer is option two. Inconsistency between near identical cases is the signature of a judgment call with no fixed criteria behind it. Naming the criteria directly removes the judgment call. It does not ask the model to somehow be more consistent about making one. Option one is the same failure as be conservative, a vibe layered on a vibe. It gives the model nothing more specific to apply than it already had. Option three treats inconsistency as random noise to average out. It is actually a missing definition. Three runs of an undefined judgment call produce three inconsistent verdicts just as easily as one. Averaging three guesses does not produce an answer. It produces an average of three guesses. Option four assumes the problem is visibility. It assumes the model would resolve the inconsistency if only it could see both files together. The pattern is genuinely ambiguous under the current instructions, regardless of what else is in view. More context does not resolve an ambiguity that was never about missing information. Three sentences on what the exam will ask. Expect a review tool with a high false positive rate. The fix is explicit criteria, plus disabling the worst offending category. Never a vaguer instruction to be careful. Expect a few shot question where the tell is inconsistent output despite clear instructions. Two to four examples showing reasoning is the answer, over one more paragraph of description. And when self review comes up, the fix is always an independent instance. Never a smarter prompt, asking the same session to catch its own mistake. Next time, the other half of structured output. An extraction pipeline that invents invoice numbers with total confidence, and why the fix is not asking it to be more careful. This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit brodynetworks.substack.com (https://brodynetworks.substack.com?utm_medium=podcast&utm_campaign=CTA_1)

4 Sept 2026
Episode 7. Plan Mode, Iteration, and CI
Domain 3, task statements 3.4, 3.5 and 3.6. A CI pipeline hangs for forty minutes waiting for a keystroke that will never come. When plan mode earns its cost and when it is just ceremony, four techniques for iterating well, and the exact flags that make Claude Code behave in an unattended pipeline. Independent and unofficial. This series is not affiliated with, sponsored by, or endorsed by Anthropic. Nothing in it is exam content. Every practice question was written for this show against the published exam guide, which is a free public document linked below. Chapters * 0:00 Cold open and disclaimer * 0:23 The pipeline that hung for forty minutes * 0:59 Plan mode: architectural weight * 1:41 Direct execution: the other legitimate choice * 2:23 The trap: size versus approaches * 2:42 Combining plan mode and direct execution * 3:13 The Explore subagent * 4:14 Four iteration techniques * 4:23 Concrete input and output examples * 4:49 Test-driven iteration * 5:29 The interview pattern * 6:02 Batching interacting fixes versus sequential * 7:03 The -p flag * 7:38 Structured CI output: json and json-schema * 8:09 CLAUDE.md as CI context * 8:36 Why an independent instance reviews better * 9:18 Giving a re-run memory of prior findings * 10:00 Worked question * 11:12 Why the other three options are there * 11:50 What the exam will ask, and next time Sources * Claude Certified Architect, Foundations Exam Guide, Version 1.0 (section 6, task statements 3.4, 3.5 and 3.6) (https://everpath-course-content.s3-accelerate.amazonaws.com/instructor%2F6nizmqk8tpzpfjvt6qmmav7rh%2Fpublic%2F1783542750%2FClaude+Certified+Architect+%E2%80%93+Foundations+Exam+Guide.pdf) * CLI reference (-p, --output-format, --json-schema) (https://docs.claude.com/en/docs/claude-code/cli-reference) * Headless and CI use (https://docs.claude.com/en/docs/claude-code/headless) Transcript This is Passing C C A R F, an independent study companion for the Claude Certified Architect, Foundations exam. Sixty items, a hundred and twenty minutes, and a published blueprint that tells you almost exactly what it is going to ask. Independent and unofficial. Not affiliated with or endorsed by Anthropic. No exam content. A pipeline sat hung for forty minutes. Nothing crashed. Nothing timed out on its own. Claude Code was simply sitting there, waiting politely for someone to type an answer to a question. Nobody was watching the machine. The job was supposed to run unattended overnight, and finish before anyone woke up. Domain three continues. Task statements three point four, three point five and three point six. Today: when to plan before you build. How to iterate well once you are building. What it actually takes to run Claude Code somewhere nobody is sitting at a keyboard. Start with plan mode. The exam tests a judgment call here, not a mechanism. Plan mode is for tasks with real architectural weight. Large scale changes. More than one valid approach on the table. Decisions that touch the shape of the system, not one function inside it. A migration touching forty five files. A choice between two infrastructure approaches with genuinely different tradeoffs. A restructuring that touches how a dozen services talk to each other. Commit to code before thinking the approach through, and you risk expensive rework. Plan mode lets that thinking happen safely, before a single line changes. Direct execution is the other half, and it is just as legitimate. Not a lesser choice. A bug fix confined to one file, where a stack trace already points at the exact line. Adding one validation check to one function. Well scoped, well understood changes. A planning phase here adds ceremony, not safety, because there was never more than one reasonable way to make the fix. Size alone will fool you here. A one line change to a shared authentication check can carry more architectural weight than a five hundred line refactor confined to one module. Count the approaches on the table, not the lines in the diff. The exam’s favorite trap is not asking you to define plan mode. It describes a task and asks whether it warrants planning at all. The tell is not the size of the diff. It is whether multiple valid approaches exist, or the path was already obvious the moment the bug was found. Real work often needs both halves, not one or the other. Plan mode for the investigation. Direct execution once the plan is settled. Work out how a library migration should actually proceed, which files it touches and in what order. Once that plan is solid, switch to direct execution and carry it out. Planning and executing are not phases you pick once and stay in. They are two tools you hand off between, inside the same piece of work. One more piece belongs here, and it solves a problem specific to investigation: context getting eaten by discovery before you have even started building. A multi phase task often needs a verbose exploration phase first. Reading a lot of files. Tracing a lot of imports. Do all of that directly in your main conversation, and your context window fills up with exploration before implementation has even started. The Explore subagent exists for exactly this. It isolates the discovery work. Only a summary comes back. Picture mapping an unfamiliar service across thirty files before touching any of them. Do that through the Explore subagent, and your main conversation gets a clean summary of what the service does and how its pieces connect. Not thirty files worth of raw exploration you will never need to see again. The context you saved is context still available for the part of the job that actually matters: writing the change itself. Now task statement three point five. Getting better output once you are actually iterating, rather than deciding whether to plan. First technique: concrete input and output examples. A prose description gets interpreted inconsistently across attempts more often than people expect. Two or three actual examples, input and correct output, communicate a transformation far more reliably than another paragraph trying to nail it down. Show the shape. Do not just describe it. Second technique: test driven iteration. Write the test suite first. Expected behavior, known edge cases, performance requirements that matter. Then iterate by sharing the actual test failures back, rather than describing the problem in your own words each time. A failing test is a precise signal. Your own summary of the problem often is not. Hand the model a stack trace and a failing assertion, and it knows exactly what wrong looks like. Hand it a paragraph of your own words, and it is reconstructing your intent from a description, one step removed from the actual failure. Third technique, and the one people underuse the most: the interview pattern. Before implementing something in an unfamiliar domain, have Claude ask you questions first, rather than diving straight into code. Cache invalidation strategy. Failure mode handling. Edge cases around concurrent access. A developer moving fast might never think to specify any of it up front. Good clarifying questions surface it before it gets baked into a wrong design, instead of after. Fourth, and this is the one with the sharpest exam signature: batching fixes versus fixing sequentially. Issues that interact, where changing one thing affects how another part of the fix should work, belong in one detailed message together. Fix them one at a time, and each fix risks undoing or complicating the next. Issues that are genuinely independent cost nothing to handle sequentially. The exam question here is rarely about the technique. It is about judging correctly whether the issues in front of you actually interact. Get that judgment wrong in either direction and it costs you. Batch independent issues together, and you have written one long, tangled message where a series of short ones would have been clearer and easier to verify one at a time. Split interacting issues apart, and the second fix quietly undoes something the first one depended on, because nobody told it the two were connected. Last task statement: getting Claude Code into a continuous integration pipeline, which is exactly where the forty minute hang I opened with happened. The fix has a name, and it is the single most testable fact in this task statement. The dash p flag, sometimes written dash dash print, runs Claude Code non interactively. No prompt ever waits for a keystroke, because there is no keystroke coming in an automated pipeline, and the flag exists specifically so the tool never tries to ask for one. Two more flags make the output usable by other software, not just readable by a human. Output format json turns the response into something your pipeline can actually parse, instead of prose you would have to scrape. Json schema goes further, enforcing that the structure matches a schema you define. That is what lets you post specific findings as inline pull request comments, instead of one giant unstructured comment block. CLAUDE.md does real work here too. It is not just for interactive sessions. In a CI context, it is how a headless invocation learns your testing standards, your fixture conventions, your review criteria, the same way it would teach any interactive session. A CI run with no CLAUDE.md context is reviewing code against no stated standards beyond generic good practice. There is a subtler point here, and it changes how you would actually architect a review pipeline. The same session that wrote a piece of code is measurably worse at reviewing it than an independent session would be. It already committed to its own reasoning while writing it. Reviewing your own work means re-examining decisions you already made confidently, and that is harder than looking at someone else’s code fresh. The consequence: your CI review step should be a separate, independent invocation, never a continuation of whatever session generated the change. Different eyes, every time, even when those eyes belong to the same model. Two more habits round this out. When a review job runs again after new commits, include the prior findings in context, and tell it to flag only what is new or still unresolved. Skip that, and every re-run repeats comments a human already read and fixed or dismissed. Do that enough nights in a row, and the team stops reading the bot at all. When generating tests specifically, provide the existing test files as context, so the generator does not propose scenarios the suite already covers. Document your testing standards and fixtures in CLAUDE.md, so generated tests are worth the tokens they cost. Let me work a question. A team runs Claude Code in a nightly CI job, reviewing the day’s merged pull requests. Configured with the dash p flag and output format json. It works reliably. After a few weeks, engineers stop reading the bot’s comments, because most nights it repeats the same three or four issues, already flagged, already fixed, in a previous run. What is the fix. Option one, remove the dash p flag so a human can interact with the review in real time. Option two, include the previous run’s findings in the current run’s context, and instruct it to report only new or still unresolved issues. Option three, switch output format json back to plain prose, so comments read more naturally. Option four, run the review weekly instead of nightly. The answer is option two. The actual defect is that each run has no memory of what a previous run already said. It re-derives the same findings from scratch, every single night. Give it that memory, and the repetition stops. Option one breaks the entire premise of unattended CI. It reintroduces the exact hang this task statement opened by warning against, to fix a problem that was never about interactivity. Option three changes how the output looks, not why the same comments keep recurring. Json versus prose is a formatting choice. It is not a memory one. Option four just makes the noise rarer, not smaller, and it delays genuinely new issues by up to a week. That trades one symptom for a worse version of the same problem. Three sentences on what the exam will ask. Expect a plan mode question that hinges on whether multiple valid approaches exist, not on how large the change sounds. Expect the dash p flag to be the answer whenever a CI job is described as hanging or waiting for input, with invented flag names as the distractors. And when repeated CI findings show up in a scenario, the fix is almost always giving the next run memory of what the last run said. Not the model. Not the schedule. Not the output format. Next time, domain four, and the gap most candidates actually have. A review bot so imprecise the team turned it off, and why the phrase be conservative did absolutely nothing to fix it. This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit brodynetworks.substack.com (https://brodynetworks.substack.com?utm_medium=podcast&utm_campaign=CTA_1)

4 Sept 2026
Episode 6. Configuring Claude Code, Memory, Rules, Commands, Skills
Domain 3, task statements 3.1, 3.2 and 3.3. A new hire’s Claude Code ignores half the team’s conventions, and the reason is not code, it is configuration scope. The three CLAUDE.md levels, splitting a sprawling file with @import and a rules directory, commands versus skills, and why a scattered convention needs a glob-scoped rule instead of a directory-level file. Independent and unofficial. This series is not affiliated with, sponsored by, or endorsed by Anthropic. Nothing in it is exam content. Every practice question was written for this show against the published exam guide, which is a free public document linked below. Chapters * 0:00 Cold open and disclaimer * 0:24 The new hire’s broken conventions * 0:44 Three configuration levels * 1:36 Why user-level settings never travel * 2:08 The fix: check which level it lives at * 2:33 The /memory command as diagnosis * 3:19 Keeping CLAUDE.md from sprawling: @import * 4:09 The rules directory, split by topic * 4:33 Commands: project versus personal scope * 4:59 Skills and their frontmatter * 6:04 Personal skill variants * 6:45 Skill or CLAUDE.md: the real judgment call * 8:13 Path-scoped rules and glob patterns * 9:50 Worked question * 10:55 Why the other three options are there * 11:56 What the exam will ask, and next time Sources * Claude Certified Architect, Foundations Exam Guide, Version 1.0 (section 6, task statements 3.1, 3.2 and 3.3) (https://everpath-course-content.s3-accelerate.amazonaws.com/instructor%2F6nizmqk8tpzpfjvt6qmmav7rh%2Fpublic%2F1783542750%2FClaude+Certified+Architect+%E2%80%93+Foundations+Exam+Guide.pdf) * CLAUDE.md hierarchy and /memory (https://docs.claude.com/en/docs/claude-code/memory) * Skills and SKILL.md frontmatter (https://docs.claude.com/en/docs/claude-code/skills) Transcript This is Passing C C A R F, an independent study companion for the Claude Certified Architect, Foundations exam. Sixty items, a hundred and twenty minutes, and a published blueprint that tells you almost exactly what it is going to ask. Independent and unofficial. Not affiliated with or endorsed by Anthropic. No exam content. A new engineer joins the team. Opens Claude Code on a repository everyone else has worked in for a year. Gets code that ignores half the team’s conventions. Claude Code is not broken. The conventions live in a file on someone else’s machine. Nobody ever put them anywhere this engineer’s copy could see. Domain three, twenty percent of the exam, tied with domain four. Task statements three point one, three point two and three point three. Where configuration lives. How commands and skills get built. How a rule finds the files it actually applies to. Start with the hierarchy, because the new hire’s problem is a hierarchy problem. Three levels. User level, in your own home directory, applies only to you. Project level, in the repository itself, applies to everyone who checks it out. Directory level, inside a specific subdirectory, applies to work happening there. A CLAUDE.md sitting inside a payments folder shapes how Claude Code behaves while working inside that folder, and nowhere else. Move to a different part of the codebase, and that file has no say. Here is the fact that explains the whole broken experience, and it catches experienced engineers too, not just newcomers. User level settings are personal. They do not travel through version control. A senior engineer can spend six months tuning their own setup and never write a line of it into the project file. All six months of tuning belongs to one person, on one machine. A new teammate inherits none of it. Not their fault. There was never anything in the repository to inherit. The fix is not a lecture about onboarding. It is a habit. When Claude Code ignores a convention you know exists, check which level it actually lives at. Sitting in someone’s personal file, it was never going to travel with the repository. That five second check saves the hour someone would otherwise spend rewriting a convention that already existed, just in the wrong place. Claude Code gives you a tool for exactly this. The slash memory command shows which memory files actually loaded this session. Run it, and you see the real list. User level file, present or absent. Project level file, present or absent. Every rule file that actually fired. Say two engineers report different behavior from the same prompt. Both think they are running the same setup. Run slash memory in each of their sessions and compare. Nine times out of ten, one of them has a personal file the other does not, and the mystery is solved in under a minute. Inconsistent behavior across sessions or teammates, that command tells you what is really in effect, not what you assume. Now, keeping a large CLAUDE.md from turning into a monster. Two mechanisms. At import lets one CLAUDE.md pull in another file. A payments package imports the standards on money handling and audit logging. A frontend package imports the standards on components and accessibility. Neither maintainer reads the other’s rules. The shared file gets written once. The alternative to at import is worse in a specific way worth naming. Paste the same standards into every package’s CLAUDE.md by hand, and the moment one standard changes, you are hunting down every copy to update it. At import means the source of truth lives in one file, and every package that references it gets the update automatically, the next time it loads. The other mechanism is a rules directory. Instead of one giant file, split conventions by topic. Testing in one file, A P I conventions in another, deployment in a third. Each file stays short enough that someone can actually find the rule they need. Nobody has to scroll past deployment conventions to find the one testing rule they actually came looking for. Task statement three point two. Commands and skills. Easy to blur if you have not built both. Commands live in a commands directory. Project scoped, shared through version control, for the whole team. Personal scoped, in your own home directory, just for you. Write a project scoped command once, and it is available to everyone who checks out the repository. Skills carry their own file, SKILL dot M D, with frontmatter worth knowing by name. Context fork comes first. A skill that produces a lot of noise, a full codebase analysis, ten brainstormed approaches, does not have to dump that into your main conversation. Run it with context fork instead. The work happens in its own sub-agent context. Only the result comes back. Allowed tools restricts what a skill can touch while it runs. A skill whose whole job is writing a report gets scoped to file writes only. Nothing else. A bug in that skill cannot reach further than the report. And argument hint prompts a developer for whatever parameter the skill needs, the moment they invoke it without one. Invoke a deploy skill with no target environment named, and instead of guessing, or failing silently, it asks. Staging or production. One clear question, instead of a skill that ran against the wrong environment because nobody told it which one. One more pattern. Want your own version of a shared skill, without touching the one everyone else uses. Give it a different name in your personal skills directory. Your version sits alongside the team’s, not in place of it. Say the team’s release skill always runs a full test suite before tagging a version. You want a faster personal variant for your own local drafts, one that skips the slow integration tests. Name it something else. Release dash draft, say. Now it lives next to the team’s release skill, not instead of it. Nobody else’s workflow changes. Nobody has to know you built a shortcut for yourself. The real judgment call in this task statement: skill, or CLAUDE.md. CLAUDE.md loads every session, automatically. A skill loads on demand. Universal standards, the ones that should shape every interaction, belong in CLAUDE.md. A specialized workflow, one only a fraction of sessions ever need, belongs in a skill. Think of it as the difference between a habit and a tool in a drawer. A habit runs every time, without being asked. A tool in a drawer only costs you anything the moment you reach for it. Load it into every session anyway, and you are paying for relevance nobody asked for. Commit message formatting goes in CLAUDE.md, since every session might eventually commit. A legacy API migration goes in a skill, since only the sessions doing that migration will ever need it. Get the judgment call backwards and both directions hurt. Put a narrow, rarely used workflow into CLAUDE.md, and every single session pays a small tax for something almost nobody needs that day. Put a universal standard into a skill instead, and it silently stops applying the moment nobody remembers to invoke it. Neither failure looks dramatic. Both just quietly cost you, one token at a time, or one missed convention at a time. Last task statement, and it has the cleanest exam signature of the three. Path specific rules. A rules file carries Y A M L frontmatter with a paths field, glob patterns, and the rule activates only when you are editing a matching file. This beats a subdirectory CLAUDE.md in one specific way, and the exam loves testing files to make the point. Say your tests are scattered across a dozen directories, not gathered into one folder. A subdirectory CLAUDE.md cannot reach all of them. There is no single directory to put it in. A glob scoped rule can. One pattern, matching every test file by name, loads the convention wherever that file happens to live. None of this makes directory level CLAUDE.md the wrong tool everywhere. A folder that genuinely contains one coherent thing, a specific microservice, a specific package, is exactly what it is built for. The failure only shows up when the convention and the folder structure do not line up. When what you are trying to govern is a file type scattered across the tree, rather than a place in it. There is a second reason this matters beyond correctness: tokens. A rule that only loads for matching files means the model is not carrying testing conventions while it works on something that has nothing to do with tests. Working on a payments function, editing zero test files, the testing rule never loads at all. No wasted context, no rule nobody is using competing for the model’s attention on that turn. Let me work a question. A team keeps testing conventions in a CLAUDE.md sitting in their tests directory. But test files live scattered across fifteen-plus directories, next to the code they test, not gathered into that one folder. Claude Code follows the conventions perfectly inside the tests directory. Everywhere else a test file actually lives, it ignores them. What is the fix. Option one, copy the same CLAUDE.md into every directory with a test file. Option two, replace the directory level CLAUDE.md with a path scoped rule, a glob matching test files by name, regardless of directory. Option three, move the CLAUDE.md to the project root. Option four, add a testing reminder to the commit message template. The answer is option two. A glob pattern reaches every test file, wherever it sits. It does not care which folder a file happens to be in. It cares what the file is. Copying is the instinct people reach for first, and it is worth understanding exactly why it fails, not just that it does. Option one works, technically, and it is the wrong shape. Fifteen copies of one file is a maintenance cost that grows every time a new test file appears somewhere else. Every copy has to be kept in sync by hand. Option three fixes coverage by brute force. Now every interaction across the entire codebase carries testing conventions, whether or not the current file has anything to do with tests. That is exactly the wasted context a glob rule exists to avoid. Reminders are not gates. A reminder gets skipped exactly as often as a prompt instruction does, and for the same reason. It reads well on paper. It has no teeth. Option four does not touch configuration at all. A line in a commit template has no mechanism for shaping how code gets written in the first place. Three sentences on what the exam will ask. Expect a new-team-member scenario, and the correct diagnosis is which configuration level a convention actually lives at, not that the convention is broken. Expect a scattered-file question where the answer is a glob scoped rule, precisely because the files are not gathered into one directory. And when skills come up, know the frontmatter by job: context fork for isolation, allowed tools for blast radius, argument hint for a missing parameter. Next time, plan mode, iteration, and getting Claude Code into a CI pipeline without it sitting there waiting for a keystroke that never comes. This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit brodynetworks.substack.com (https://brodynetworks.substack.com?utm_medium=podcast&utm_campaign=CTA_1)

4 Sept 2026
Episode 5. MCP in Production, Errors, Scoping, Built Ins
Domain 2, task statements 2.2, 2.4 and 2.5. An agent stuck retrying a permission error against a tool that only ever says operation failed. The four MCP error classes, project versus user scoped server configuration, MCP resources as content catalogs, and picking the right built-in tool for the job instead of the familiar one. Independent and unofficial. This series is not affiliated with, sponsored by, or endorsed by Anthropic. Nothing in it is exam content. Every practice question was written for this show against the published exam guide, which is a free public document linked below. Chapters * 0:00 Cold open and disclaimer * 0:23 An agent retrying a locked door * 0:52 The isError flag, and its limit * 1:37 Four error classes * 2:17 Structured metadata: category, retryable, description * 3:18 Local recovery versus propagating upward * 4:11 Access failure versus a valid empty result * 5:07 Project versus user scoped MCP config * 6:02 Credentials via environment variable expansion * 6:33 Writing MCP descriptions the agent will prefer * 7:13 Community servers over custom ones * 7:37 MCP resources as content catalogs * 8:07 Grep, Glob, Read, Write, Edit * 8:43 The Read plus Write fallback * 9:16 Exploring an unfamiliar codebase * 9:57 Worked question * 11:00 Why the other three options are there * 11:51 What the exam will ask, and next time Sources * Claude Certified Architect, Foundations Exam Guide, Version 1.0 (section 6, task statements 2.2, 2.4 and 2.5) (https://everpath-course-content.s3-accelerate.amazonaws.com/instructor%2F6nizmqk8tpzpfjvt6qmmav7rh%2Fpublic%2F1783542750%2FClaude+Certified+Architect+%E2%80%93+Foundations+Exam+Guide.pdf) * MCP in Claude Code (https://docs.claude.com/en/docs/claude-code/mcp) Transcript This is Passing C C A R F, an independent study companion for the Claude Certified Architect, Foundations exam. Sixty items, a hundred and twenty minutes, and a published blueprint that tells you almost exactly what it is going to ask. Independent and unofficial. Not affiliated with or endorsed by Anthropic. No exam content. The agent hit a permission error on a backend call, and it tried again. Then it tried again. Four more times, each with a slightly different approach, each one failing for the exact same reason, because the tool’s response, every time, was two words: operation failed. Nothing about why. Nothing about whether trying again could possibly help. The agent had no way to know it was banging on a locked door, so it kept knocking. Domain two continues. Task statements two point two, two point four and two point five. Error responses, server scoping, and the built in tools everyone already has and half the time reaches for wrong. Start with the error, because it is the cleanest version of a pattern you have already seen in this series: a uniform response hides information the agent needs to act well. The Model Context Protocol has a mechanism for this, an is error flag that marks a tool result as a failure rather than a success. That flag alone tells the model something went wrong. It does not tell the model what kind of wrong, and that distinction is the entire episode. There are four classes of error worth telling apart, and the exam wants you fluent in all four. Transient errors, a timeout, a service that is briefly unavailable, where trying again in a moment might genuinely work. Validation errors, the input itself was malformed, where retrying the identical call will fail identically forever. Business errors, the operation was understood perfectly and refused on purpose, a policy violation, not a glitch. And permission errors, which is exactly what opened this episode, where the caller is not allowed to do this at all, and no number of retries changes that. Notice what those four categories buy you that a flat operation failed cannot. They tell the agent whether retrying is even worth attempting. A generic failure response leaves that question unanswered. So the agent either gives up too early on something recoverable, or burns time retrying something that was never going to succeed no matter how many times it asked. The fix is structured metadata, and the shape of it is worth knowing by name. An error category field, transient, validation, business or permission. An is retryable boolean, so the agent does not have to guess from the category alone. And a human readable description that can be surfaced to an actual customer when the failure is the kind of thing a customer needs to hear about. For business errors specifically, that readable explanation matters even more. A policy violation is not a bug to route around. It is information the person on the other end deserves to receive in plain language. Here is where the four category, is retryable pattern actually saves you real cost, and it is the same logic as the enforcement episode a few back. Local recovery belongs inside the subagent that hit the failure. If a call times out and the category is transient, retry it there, quietly, without ever bothering the coordinator. What should travel upward to the coordinator is only the failure that could not be resolved locally. It should travel with two things attached: what was actually attempted, and whatever partial results already exist. A coordinator that receives a bare failure notice with no partial results has to start that piece of the investigation from nothing. A coordinator that receives a structured failure with partial results attached can often keep going with what it already has. One more distinction inside this same territory, and it is subtle enough that people genuinely mix it up under time pressure. An access failure and a valid empty result are not the same thing, even though both can look, from a distance, like nothing came back. An access failure means the query could not run, permission denied, service unreachable, something is actually broken. A valid empty result means the query ran perfectly and there is nothing in your data matching that particular ask. Treating an empty result as a failure means retrying a search that was never going to return anything different. Treating a real failure as an empty result means quietly reporting success on a query that never actually executed. Both are wrong in opposite directions, and the fix for both is the same: know which one you are looking at before you decide what to do next. Now scoping, task statement two point four, and this is mostly about where configuration lives rather than what it says. Two levels. Project scoped configuration, in dot m c p dot Jason, is for shared team tooling, the servers everyone on the project needs and that travel with the repository in version control. User scoped configuration, in your home directory’s claude dot Jason, is for personal or experimental servers. Things you are trying out, that have no business being forced on every other person who checks out the repo. Both levels are active at once. Tools from every configured server, project and user, are discovered when the connection is made and are all available to the agent simultaneously. This is not an either or choice, it is two pools that both feed the same agent. Credentials belong in the project file without ever actually being in the project file, and the mechanism is environment variable expansion. Write a token as a reference to an environment variable rather than the literal secret, and the actual value is pulled from the environment at run time. This is how a team shares MCP server configuration in git without also sharing every team member’s API keys in git, which would otherwise be the obvious and disastrous alternative. There is a quieter skill buried in task statement two point four that is easy to skip past: writing MCP tool descriptions well enough that the agent actually prefers them over a built in. Say you connect a capable MCP tool for ticket lookups, and its description is thin. The agent may keep reaching for a generic built in search instead. The built in tool is the path of least resistance, and the MCP tool never explained why it is the better choice. The fix is the same one from two episodes back: write the description like it is the interface, because it is. And a related judgment call: reach for an existing community MCP server before building your own, for anything standard. Most teams do not need a custom Jira integration when a maintained community one already does the job. Save custom servers for the workflows that are genuinely specific to your team, where nothing off the shelf fits. Last idea in this task statement, and it is a nice one: resources. A resource is not a tool you call to take an action. It is a catalog you can browse, exposed by the server itself: a list of issue summaries, a documentation hierarchy, a database schema. Giving an agent access to that catalog directly means it does not need a round of exploratory tool calls first. It can see what exists before it decides what to actually ask for. Finally, task statement two point five, the built in tools you already have without connecting anything. Grep searches file contents, the tool for finding a function name, an error message, or every place a particular string appears across a codebase. Glob matches file paths by pattern, the tool for finding files by name or extension rather than by what is written inside them. Read and Write handle whole files. Edit makes a targeted change by matching unique surrounding text, and it is precise exactly because it insists on uniqueness. That insistence is also where Edit sometimes fails, and the fallback matters because it comes up constantly in real work. If the text you are trying to anchor an edit to appears more than once in a file, Edit cannot tell which occurrence you mean, and it will refuse rather than guess wrong. The reliable fallback in that situation is Read the whole file, make the change yourself in memory, then Write the full file back. Not a workaround, a documented fallback for exactly this case. There is also a way of working through an unfamiliar codebase that the exam rewards over the more obvious approach. Do not start by reading every file. Start with Grep to find entry points, the handful of places a feature actually begins. Then use Read to follow the imports and trace the flow from there. Expand only into the files that matter for the question you are actually asking. Tracing usage of a function across wrapper modules follows the same shape. First find every exported name involved. Then Grep for each of those names across the codebase, rather than opening every file and hoping the connection jumps out at you. Let me work a question. An MCP tool that looks up account balances returns operation failed for three different underlying situations. The account ID was submitted in a malformed format, letters where only digits are valid. The caller lacks permission to view the account. Or the backend service is temporarily down. An agent using this tool retries every operation failed the same way, three times, with a short delay between attempts. What is the most direct fix. Option one, increase the retry count and the delay between attempts. Option two, have the tool return structured error metadata, an error category of validation, permission, or transient depending on the actual cause, with an is retryable flag. Option three, wrap the tool in a subagent dedicated entirely to balance lookups. Option four, add a system prompt instruction telling the agent when it is appropriate to retry. The answer is option two. The tool is currently discarding the exact information that would let the agent behave differently across these three cases, and returning that information is what actually fixes the behavior rather than tuning around it. Option one makes every case, including the two that can never succeed by retrying, take longer to fail, without changing whether they fail. Option three adds structure without addressing the return value the new subagent would still receive, so the same three collapsed cases arrive at it unchanged. Option four asks a prompt to compensate for missing data the model has no way to reconstruct. There is no cue in a bare operation failed message that lets the agent infer which of three situations it is looking at, no matter how well the instruction is worded. Three sentences on what the exam will ask. Expect a scenario where an agent mishandles retries because an error response gave it nothing to distinguish transient from permanent, and the fix is structured error metadata, not retry tuning. Expect a scoping question that tests whether you know project configuration is for shared team servers and user configuration is for personal ones, both active at once. And when Grep, Glob, Read, Write and Edit show up together, the exam is usually testing whether you pick the narrowest tool for the actual operation, not whether you know they all exist. Next time, domain three. A new hire joins the team, and Claude Code ignores every convention the codebase actually follows, because nobody thought about where a convention lives before writing it down. This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit brodynetworks.substack.com (https://brodynetworks.substack.com?utm_medium=podcast&utm_campaign=CTA_1)

4 Sept 2026
Episode 4. Tool Descriptions Are the Interface
Domain 2, task statements 2.1 and 2.3. Two tools with nearly identical descriptions, and a model that keeps confusing them. What actually belongs in a tool description, why a well written system prompt can quietly override one, and how many tools is too many for reliable selection. Independent and unofficial. This series is not affiliated with, sponsored by, or endorsed by Anthropic. Nothing in it is exam content. Every practice question was written for this show against the published exam guide, which is a free public document linked below. Chapters * 0:00 Cold open and disclaimer * 0:21 Two tools that keep getting confused * 0:48 Why a description is the interface * 1:39 What goes into a description that routes * 3:17 Fixing overlap: rename, or split * 4:15 System prompt keywords that override descriptions * 5:29 Eighteen tools versus four or five * 6:10 Tools outside a role, and misuse * 6:48 Scoped cross-role exceptions * 7:13 Replacing a generic tool with a constrained one * 7:44 auto, any, and forced selection * 9:09 Worked question * 10:07 Why the other three options are there * 11:19 What the exam will ask, and next time Sources * Claude Certified Architect, Foundations Exam Guide, Version 1.0 (section 6, task statements 2.1 and 2.3) (https://everpath-course-content.s3-accelerate.amazonaws.com/instructor%2F6nizmqk8tpzpfjvt6qmmav7rh%2Fpublic%2F1783542750%2FClaude+Certified+Architect+%E2%80%93+Foundations+Exam+Guide.pdf) * Tool use overview (tool_choice options) (https://docs.claude.com/en/docs/agents-and-tools/tool-use/overview) Transcript This is Passing C C A R F, an independent study companion for the Claude Certified Architect, Foundations exam. Sixty items, a hundred and twenty minutes, and a published blueprint that tells you almost exactly what it is going to ask. Independent and unofficial. Not affiliated with or endorsed by Anthropic. No exam content. Two tools sat next to each other in the same agent. One was called analyze content. The other was called analyze document. Their descriptions were nearly identical, a sentence each, competent enough on their own, meaningless side by side. The model routed customer questions to whichever one it felt like that turn, and about a third of the time it felt like the wrong one. Nobody had touched the underlying logic. Nobody needed to. Domain two, tool design and Model Context Protocol integration, eighteen percent of the exam. Task statements two point one and two point three today. The subject is smaller than it sounds and more consequential than people expect: what a tool description actually is, and how many tools is too many. Here is the idea underneath the whole episode. A tool description is not documentation for a human reading the code later. It is the interface the model uses to decide which tool to call, in real time. The only information it has about your intent is the words you wrote in that description, plus the words in your prompt. Write it thin, and the model is guessing. Write it well, and routing mostly takes care of itself. A minimal description gets you unreliable selection the moment two tools do anything similar. This is exactly what happened to analyze content and analyze document. Neither description said what kind of content, what kind of document, what one could do that the other could not, or when a human building the system would reach for one over the other. The model had no boundary to use, so it stopped using one reliably. What actually goes into a description that routes correctly. Input formats, so the model knows what shape of data belongs here. Example queries, so it can pattern match a real request against a real use case rather than guessing from a name. Edge cases, so it knows what this tool is not for. And boundaries against the similar tool sitting right next to it, spelled out directly rather than left for the model to infer from two vague sentences. Picture what that actually looks like written out. Not “looks up order information,” which is the kind of sentence that started this episode’s problem. Something closer to this. It accepts an order ID or a customer email. It returns shipment status and line items. Use it when the customer references a specific order. It does not handle billing disputes or account changes, which live in separate tools. That last clause, the explicit does not handle, is doing as much work as everything before it. It is drawing the boundary the model needs in order to know when a different tool is the right call instead of this one. The fix for analyze content and analyze document was not a smarter prompt around them. It was renaming and rewriting. Analyze content became extract web results, with a description specific to pulling structured data out of web pages. The overlap did not get patched, it got removed, because the two tools no longer described anything similar enough to confuse. Sometimes the right move is not renaming one tool, it is splitting one tool into several. Take a generic document tool that tries to do everything: pull data points, summarize, check a claim against a source. Its description ends up either too vague to route well, or too long to read cleanly. Split it into three narrower tools instead, one per job, each with its own tight input and output contract. Each one becomes easy to describe precisely, because it only does one thing. There is a failure mode here that has nothing to do with the tool description itself, and it is the one people miss. System prompts contain instructions, and instructions contain keywords, and keywords can quietly override a perfectly well written tool description. Say your system prompt says something like always check the latest data before answering. If one tool’s description mentions the word latest, the model may reach for that tool constantly, whether or not it is the right one. The system prompt just handed it an unintended association. Reviewing the system prompt for this kind of keyword sensitivity is part of the actual skill here, not an afterthought. A tool description problem sometimes lives one document away from the tool. Worth sitting with for a second, because it explains why a team can rewrite a tool description three times and still see the same misrouting. If the bug is not in the description at all, no amount of rewriting the description touches it. The actual fix is reading the system prompt with the same scrutiny you would give the tool descriptions, looking for any word that happens to also appear in a tool’s name or summary. Now the second half of the episode, and it is a different kind of overload. Task statement two point three: how tools get distributed across agents, and how you tell the model which one to use when more than one is available. The guide’s own contrast is worth remembering exactly, because it is concrete enough to stick. Eighteen tools versus four or five. Give an agent eighteen tools and selection reliability degrades, not because any individual tool is badly described, but because decision complexity itself has a cost. More options to weigh means more chances to weigh them wrong, even when every option is well documented. This shows up as a specific, predictable failure. Hand an agent a tool outside its own specialty, and sooner or later it reaches for that tool when it should not. Say a synthesis agent has been handed a web search tool, because it seemed convenient to give it access to everything. It will sometimes go searching the web mid-synthesis, instead of doing the one job it was built for. The presence of the tool creates the temptation. The fix is scoped access: give each agent only the tools its role actually needs, and resist the instinct to hand out extra tools just in case. That does not mean zero cross-role tools, ever. Some needs are frequent enough across roles that a narrow, scoped exception earns its place. A narrow fact-checking tool, scoped to the synthesis agent specifically, for a quick check it needs constantly. Anything more complex still routes back through the coordinator, rather than being solved by handing every agent every tool. Sometimes distribution is not the fix, replacement is. A generic fetch url tool available broadly is an invitation for an agent to fetch something it should not, or to fetch a document in a shape downstream code cannot handle. Replacing it with something narrower, load document, that validates the URL and the document type before returning anything, removes an entire category of misuse by construction rather than by hoping the model behaves. Last piece, and it is the one exam candidates most often blur together: tool choice configuration. Three settings, three different jobs, and each one earns its place in a different kind of turn. Auto lets the model decide whether to call a tool at all. This is the right default for most conversational work, a general assistant that sometimes needs a tool and sometimes just needs to answer in plain language. Any forces the model to call some tool, whichever one it judges best, rather than returning plain text. Reach for this when a chatty non-answer would actually be a failure. A classification agent whose whole job is producing a structured verdict on every single input should never be allowed to reply with an explanation instead of a call. Forced selection goes one step further and names one specific tool directly, guaranteeing that exact tool runs. This is what a pipeline needs when step one has to happen before step two can even make sense. Force the extraction step first, before any enrichment tool is even offered as an option, and handle the rest of the pipeline in follow-up turns, once that first result exists. Auto would leave that ordering to chance. Any would guarantee a tool call, but not which one. Only forced selection guarantees the specific one this step of the pipeline actually needs. Let me work a question. An agent has six tools available: three for its actual specialty, and three general purpose utilities included because the team figured they might be useful someday. Over time, the team notices the agent occasionally calls one of the general utilities in situations where a specialty tool was clearly the better fit, producing a technically valid but lower quality result. What is the most direct fix. Option one, rewrite the specialty tools’ descriptions to be even more detailed than they already are. Option two, remove the three general purpose utilities from this agent’s available tools, since they sit outside its specialization. Option three, set tool choice to forced selection on one of the specialty tools. Option four, add a system prompt instruction telling the agent to prefer its specialty tools. The answer is option two. The specialty tools are already described well enough that they are the better fit here. The general utilities are not badly written either. They are simply present when they should not be. The direct fix is scoping the tool list itself. Take away the temptation, and there is nothing left to reach for by mistake. That is worth more than any amount of extra wording. A tool that cannot be reached cannot be misused, and that guarantee costs nothing to maintain going forward. Option one adds effort to descriptions that were not the actual problem, since the agent is choosing correctly when it has no distraction available. Option three is too blunt, forced selection locks in one specific tool for every call, which breaks the cases where a different specialty tool, or no tool at all, was genuinely the right choice. Option four treats a distribution problem as a wording problem, and it is the same class of mistake as trying to fix the identity verification skip with a stronger sentence. An instruction is competing with the tool’s own presence, and removing the tool wins that argument permanently instead of most of the time. Three sentences on what the exam will ask. Expect two tools with thin, overlapping descriptions and a model confusing them, and the answer is expanding or splitting the descriptions, not merging the tools or adding a routing layer. Expect a tool count question where the correct move is narrowing an agent’s tool list to its specialization rather than writing better prompts around a crowded one. And know the three tool choice settings by what job each one does, since the exam will describe a scenario and expect you to name the setting rather than define it from scratch. Next time, domain two continues. What a Model Context Protocol tool should actually say when it fails, and why an agent that keeps retrying a permission error is a symptom of an error response that told it nothing. This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit brodynetworks.substack.com (https://brodynetworks.substack.com?utm_medium=podcast&utm_campaign=CTA_1)

4 Sept 2026
Episode 3. Hard Guarantees, Hooks, Gates, and Sessions
Domain 1, task statements 1.4, 1.5 and 1.7. A refund agent that skips identity verification twelve percent of the time, and why the fix is never a stronger sentence. When a business rule needs a programmatic gate instead of a prompt, the two shapes a hook can take, structured handoffs for human escalation, and choosing between resuming a session and starting fresh with a summary. Independent and unofficial. This series is not affiliated with, sponsored by, or endorsed by Anthropic. Nothing in it is exam content. Every practice question was written for this show against the published exam guide, which is a free public document linked below. Chapters * 0:00 Cold open and disclaimer * 0:24 The twelve percent * 1:07 Why a prompt instruction is probabilistic * 1:44 A case where a gate would be overkill * 2:26 Programmatic enforcement: prerequisite gates * 2:59 The question that decides which you need * 3:56 Two hook shapes * 4:08 Incoming: normalizing data with PostToolUse * 4:45 Outgoing: interception hooks and policy * 5:16 What both hook shapes have in common * 5:59 Multi-concern messages and shared context * 7:06 Structured handoffs for human escalation * 7:52 Session management: two mechanisms * 8:13 fork_session and the testing-strategy example * 8:46 Resume, or start fresh with a summary * 9:40 Worked question * 10:37 Why the other three options are there * 11:43 What the exam will ask, and next time Sources * Claude Certified Architect, Foundations Exam Guide, Version 1.0 (section 6, task statements 1.4, 1.5 and 1.7) (https://everpath-course-content.s3-accelerate.amazonaws.com/instructor%2F6nizmqk8tpzpfjvt6qmmav7rh%2Fpublic%2F1783542750%2FClaude+Certified+Architect+%E2%80%93+Foundations+Exam+Guide.pdf) * Claude Code hooks (https://docs.claude.com/en/docs/claude-code/hooks) Transcript This is Passing C C A R F, an independent study companion for the Claude Certified Architect, Foundations exam. Sixty items, a hundred and twenty minutes, and a published blueprint that tells you almost exactly what it is going to ask. Independent and unofficial. Not affiliated with or endorsed by Anthropic. No exam content. Twelve percent of the time, the refund agent skipped identity verification and processed the refund anyway. Not because the instructions were unclear. The system prompt said, plainly, verify identity before processing any refund. Eighty eight percent of the time, that sentence worked. Twelve percent of the time, it did not, and the difference between those two numbers is most of this episode. Domain one, task statements one point four, one point five and one point seven. Enforcement, hooks, and session management. The theme underneath all three is the same question asked in different places: when is asking nicely enough, and when do you need a guarantee. Start with the twelve percent, because it is the cleanest version of the idea. A prompt instruction is probabilistic. It shapes what the model is likely to do, and a well written one shapes it strongly, but likely is not the same word as always. Somewhere in that eighty eight to twelve split, the model decided the situation in front of it did not obviously require the step you told it never to skip, and it moved on. That is not a broken model. That is what asking a language model to enforce a rule through prose will always cost you, some rate above zero. Contrast that with a case where a gate would be overkill. An internal tool that drafts release notes from a changelog benefits from a prompt instruction reminding it to check the version number before writing the summary. If it forgets once in a hundred drafts, a human reviewing the draft catches it in five seconds, and nothing downstream is harmed. Building a programmatic gate around that step would be effort spent defending against a mistake that costs almost nothing to catch. The two situations are not different because one prompt is better written than the other. They are different because one failure is expensive and the other is not. Programmatic enforcement is different in kind, not just in strength. A hook, or a prerequisite gate, sits in code, outside the model’s judgment entirely. It blocks a downstream tool call until an upstream one has actually returned what it needs to return. The refund tool simply cannot be called until the verification tool has returned a confirmed identity. There is no version of this where the model decides the step is unnecessary this time, because the model’s opinion was never part of the mechanism. Here is the question that tells you which one you need, and it is worth carrying around as a habit rather than an exam fact. Does a failure here have financial or safety consequences. If yes, you want a gate, not a better sentence. If the downside of a miss is mild, a well written prompt with good instructions is proportionate, and building a hook for every possible workflow step is its own kind of waste. The exam rewards recognizing which situation you are in, and the identity verification scenario is written specifically so that enhanced prompting and few shot examples look like reasonable answers. They are reasonable, for a lower stakes problem. They are the wrong answer here, because they are still probabilistic, and probabilistic is exactly what a financial operation cannot afford. Hooks come in two shapes worth telling apart, and they point in opposite directions through the pipeline. One shape normalizes data coming in. The other blocks calls going out. Take the incoming shape first. Say you have three backend systems, and each reports timestamps differently. One gives you a Unix epoch number, one gives you an ISO date string, and one gives you a numeric status code that means something different in each system’s own documentation. A PostToolUse hook intercepts a tool’s result before the model ever sees it, and normalizes it into one consistent shape. The model then reasons over clean, uniform data every time, instead of relearning three different formats every single call and occasionally getting one of them wrong. The outgoing shape is the interception hook, and this is the one that actually enforces a business rule rather than just tidying data. It sits between the model deciding to call a tool and that tool actually running, checks the call against a policy, and blocks it if the policy is violated. Refunds above five hundred dollars redirect to a human escalation workflow instead of executing. The model can propose whatever it wants. The hook decides what is actually allowed to happen. Notice what both hook shapes have in common. Neither one is a suggestion. A PostToolUse hook that normalizes data runs on every single result, unconditionally. An interception hook that blocks a policy violation blocks it every time the condition is met, not most of the time. That is the property prompt instructions cannot offer you, and it is the entire reason hooks exist as a separate mechanism instead of just being better prompt engineering. A prompt that says always do this still has to be read, weighed, and followed by the model on every single turn. A hook has nothing to weigh. It runs, or the call it is watching does not happen. One more piece of task statement one point four, and it belongs right here because it is also about structure rather than wording. Real customer messages rarely carry one clean request. A single message might raise a billing question, a shipping complaint, and a request to update an address, all in the same three sentences. Handling that well means decomposing the message into its distinct concerns first, investigating each one, and only then synthesizing a single reply. Investigate them using the same shared context, rather than starting three unrelated conversations. Where the concerns do not depend on each other, there is no reason to handle them one after another instead of together. What you do not want is an agent that reads a multi-concern message, latches onto the first thing it recognizes, answers only that, and calls the turn finished. That is a coverage failure with the same shape as the narrow decomposition bug from two episodes back, just at the level of a single message instead of a whole research task. Now the handoff problem, which sits right next to enforcement because it shows up at the same moment, the moment something needs a human. When an agent escalates mid process, whoever picks it up on the human side cannot see the conversation transcript. They see whatever you hand them, and if that handoff is a vague note, they start from zero on a problem the agent already partly solved. A structured handoff includes the customer’s identifying details, a stated root cause rather than just a symptom, and a recommended action, not just a description of what went wrong. The difference between a good handoff and a bad one is the difference between a human picking up mid-investigation and a human starting the investigation over. Last topic, and it is a quieter one: session management. Two mechanisms, and they solve different problems. Named session resumption lets you continue a specific prior conversation by name. That matters once you are running more than one investigation at a time, and need to come back to a particular one later without losing its history. Fork session takes a different shape. It branches off from a shared analysis baseline into two or more independent paths. That is what you want when you have done common groundwork, and now need to explore two different approaches from that same starting point without one contaminating the other. Comparing two different testing strategies against the same codebase analysis is exactly this shape: one shared baseline, two forked branches, neither one’s exploration bleeding into the other’s results. The judgment call the exam actually tests here is not which mechanism exists, it is when to resume and when to start fresh with a summary instead. Resuming is right when the prior context is still valid, the files you looked at have not changed, the facts you gathered still hold. Starting fresh with an injected summary is right when the underlying reality has moved since that session ran. Say code has been edited since a session last looked at it. Resuming that session and asking it to keep working means it reasons from tool results that are now stale, possibly confidently wrong about what a file currently contains. The fix in that case is not resuming and hoping the model notices the drift on its own. It is telling the resumed session explicitly what changed, so it knows to re-check rather than trust its memory. Let me work a question. A team’s customer support agent has a system prompt instructing it to check a customer’s account status before offering any discount, worded clearly, with an example. Over a large sample of conversations, it skips that check and offers a discount anyway in roughly one in fifteen cases, more often when the customer’s message is emotionally charged. What is the correct fix. Option one, rewrite the instruction with stronger language and add three few shot examples of the check being performed correctly. Option two, add a programmatic prerequisite gate that blocks the discount tool from being called until the account status tool has returned a result in that conversation. Option three, lower the model’s temperature setting to make its behavior more consistent. Option four, add a PostToolUse hook that logs every discount offered for later audit. The answer is option two. An emotionally charged message correlates with the skip rate, which is exactly the signature of a prompt instruction losing a probabilistic contest against another pressure in the conversation. A gate removes that contest entirely, since the discount tool has no path to running without the prerequisite already having returned. Option one improves the odds without removing the failure mode, and the correlation with emotional tone is a strong hint that a purely textual instruction is already competing and sometimes losing. Option three is a real lever in general, but it is not a documented enforcement mechanism for this kind of business rule, and a lower temperature reduces variance without guaranteeing compliance. Option four is the most interesting wrong answer, because logging is genuinely valuable and belongs in a mature system. It only tells you afterward that the skip happened. It does not stop the fifteenth conversation from happening the way the first fourteen did. Three sentences on what the exam will ask. Expect a scenario with a measured failure rate on a rule with financial or safety stakes, and the correct answer is almost always a programmatic gate rather than a stronger prompt. Expect a hook question that asks you to tell a normalizing PostToolUse hook apart from a blocking interception hook by what each one actually does to the data flow. And on session questions, the tell is whether the world changed since the session last looked at it, not whether the mechanism name sounds more advanced. Next time, domain two, and the interface most systems get wrong first. Two tools with almost identical descriptions, and a model that keeps confusing them. We work out why the fix is never a bigger model, and it is usually not a routing layer either. This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit brodynetworks.substack.com (https://brodynetworks.substack.com?utm_medium=podcast&utm_campaign=CTA_1)

4 Sept 2026
Episode 2. Coordinators and Subagents
Domain 1, task statements 1.2 and 1.3. A research report comes back well cited and polished, and missing an entire half of its subject, while every subagent that ran reports success. The episode builds the hub-and-spoke pattern from the ground up: why routing everything through a coordinator buys observability, the isolation rule and the bug it causes, and the mechanics, Task tool, AgentDefinition, parallel spawning, that the exam tests directly. Independent and unofficial. This series is not affiliated with, sponsored by, or endorsed by Anthropic. Nothing in it is exam content. Every practice question was written for this show against the published exam guide, which is a free public document linked below. Chapters * 0:00 Cold open and disclaimer * 0:22 The report that missed half its subject * 0:50 Hub and spoke, and why it buys observability * 1:53 What the coordinator actually does * 2:22 The isolation rule * 3:20 The bug this scenario is built from * 4:06 Fixing it: evaluate, re-delegate, re-invoke * 4:42 Partitioning to avoid duplication * 5:16 The Task tool and allowedTools * 5:49 AgentDefinition: description, prompt, restrictions * 6:23 Passing complete findings explicitly * 6:59 Separating content from metadata * 7:45 Parallel spawning in one response * 8:19 Goals and criteria over fixed procedure * 9:07 Worked question * 10:20 Why the other three options are there * 11:34 What the exam will ask, and next time Sources * Claude Certified Architect, Foundations Exam Guide, Version 1.0 (section 6, task statements 1.2 and 1.3) (https://everpath-course-content.s3-accelerate.amazonaws.com/instructor%2F6nizmqk8tpzpfjvt6qmmav7rh%2Fpublic%2F1783542750%2FClaude+Certified+Architect+%E2%80%93+Foundations+Exam+Guide.pdf) * Agent SDK subagents and AgentDefinition (https://docs.claude.com/en/api/agent-sdk/subagents) Transcript This is Passing C C A R F, an independent study companion for the Claude Certified Architect, Foundations exam. Sixty items, a hundred and twenty minutes, and a published blueprint that tells you almost exactly what it is going to ask. Independent and unofficial. Not affiliated with or endorsed by Anthropic. No exam content. The report came back beautifully written. Every claim cited, every source dated, four subagents reporting success. It covered a century of a topic in careful, confident prose. It was also missing an entire half of that century’s subject matter, because nobody had told any of the four agents that half existed. Every downstream agent had done exactly what it was asked. The bug was upstream of all of them. This is domain one again, task statements one point two and one point three, orchestration and subagent configuration. Twenty seven percent of the exam, and this is the episode where the pattern most people picture when they hear the word agent actually gets built. Start with the shape. A coordinator agent sits in the middle. Subagents sit around it, each doing one kind of work. All communication runs through the coordinator. A search subagent never talks to a synthesis subagent directly. It reports back to the coordinator, and the coordinator decides what happens next. This is called hub and spoke, and the reason to build it this way rather than letting subagents talk freely is not aesthetic. It is observability. Every piece of information that moves through the system passes through one place you can log, inspect and control. A tangle of subagents messaging each other directly is a system where a failure can hide in an edge you never thought to check. The coordinator does four things. It decomposes the task into pieces. It delegates each piece to the right subagent. It aggregates what comes back. And it decides, based on how complex the query actually is, which subagents get invoked at all. That last part matters more than it sounds. A coordinator that always runs the full pipeline regardless of what was asked is not really coordinating, it is just a fixed sequence with extra steps. Now the rule that causes the failure I opened with. Subagents run in their own isolated context. A subagent does not automatically pick up the coordinator’s prior conversation, and it shares no memory with any other subagent between invocations. Each one starts fresh. It knows only what you put in its prompt. Sit with that, because it is the single most common bug in this pattern, and the exam knows it. If you write a coordinator that says “have the synthesis agent write up the findings” without including what those findings actually were, the synthesis agent has nothing. It cannot see what the search agent saw. It cannot see what the document analysis agent found. It only sees the words you gave it in that one prompt. If you did not put the findings there, they do not exist as far as that agent is concerned. This is exactly what happened in the report I opened with. The coordinator decomposed a broad historical topic into scope for its subagents, and it decomposed too narrowly, missing whole categories of the subject entirely. Every subagent it did invoke succeeded completely. There was no error anywhere in the transcript. The report was accurate and well cited about the part of the topic anyone had thought to ask about. The exam will build exactly this kind of question, and the trap is that your eye goes to the subagent that produced the final report, since that is where the gap is visible. The actual defect is upstream of it, in how the coordinator split the work in the first place. The fix at the coordinator level is to design for coverage, not just for delegation. When a topic is broad, the coordinator should check its own decomposition for holes before it ever spawns a subagent. After synthesis comes back, it should evaluate that output for gaps, and be willing to re-delegate with more targeted queries if something is missing. That loop, evaluate, re-delegate, re-invoke synthesis, repeat until the coverage is actually there, is part of the published objective. A coordinator that spawns once and reports whatever comes back is doing half the job. Partitioning matters too, and it runs the other direction from the failure I opened with. Where narrow decomposition misses things, sloppy decomposition duplicates them. If two subagents are handed overlapping scope, they research the same ground twice and burn the same time twice. The coordinator ends up reconciling two versions of the same finding, instead of getting two different findings. Assign distinct subtopics, or distinct source types, to each subagent, so the pieces actually tile instead of overlapping. Now the mechanics, because the exam tests these specifically and they are easy to get backwards. Subagents are spawned through a tool, and Anthropic calls it the Task tool. If you are building a coordinator, its list of allowed tools has to include Task, or it has no way to spawn anything. This shows up on the exam as a configuration question dressed up as a bug report: a coordinator that never delegates. The fix is not a smarter prompt. It is checking whether Task is even in the tool list. Each subagent type is defined by what Anthropic calls an AgentDefinition. That configuration is where you set the subagent’s description, so the coordinator knows when to use it. It is also where you set the subagent’s own system prompt, which shapes how it behaves once invoked, and its tool restrictions, which limit what it can touch. A search subagent probably should not have write access to your production database. Scoping that at the definition level means you are not relying on the subagent to police itself. Because subagents do not inherit context, passing it explicitly is not optional, it is the whole mechanism. The skill here is including complete findings directly in the next subagent’s prompt. If the search subagent found six sources, and the synthesis subagent needs to write about them, put all six directly into the synthesis prompt. Do not summarize them down to save space and then wonder why synthesis produced something vague. And do not assume that because the coordinator saw the search results, everyone downstream saw them too. There is a specific way to pass that content that the exam cares about, and it is a formatting decision more than a writing one. Separate content from metadata. Say a finding has a claim, a source, a page number and a publication date. Do not fold all of that into one paragraph of prose and hope the next agent parses it back out correctly. Use a structured format where the claim is one field and the source is another. This is what lets attribution survive being passed hand to hand through three or four agents. Lose that structure early and by the time a report gets written, claims and sources have drifted apart, or merged, or vanished, and nobody can tell which source said what. One more mechanical point, and it is where a lot of systems leave real speed on the table without realizing it. Say you need three subagents to run, and their work does not depend on each other. Spawn all three by emitting multiple Task calls in a single coordinator response. Do not send one call per turn, waiting for each to finish before starting the next. Parallel spawning like this is how a multi-agent system actually gets faster than doing the work yourself, instead of just being a more expensive way to do it sequentially. Last piece for this episode, and it is a design instinct rather than a mechanism. Coordinator prompts should specify goals and quality criteria, not step by step procedure. Tell a research coordinator what a complete answer needs to cover, and what would count as insufficient, and let it decide how to get there. Write it as a procedure instead, a fixed script of exactly which subagent runs when, and you have built something brittle. It cannot adapt when a search comes back thin, or a topic turns out to need a different kind of source than you expected. The goal driven version is what lets the evaluate, re-delegate loop from earlier actually function, because the coordinator has room to decide re-delegation is warranted in the first place. Let me work a question. A team builds a research coordinator. It spawns a web search subagent and a document analysis subagent for every query, then always passes their combined output to a synthesis subagent, which formats the final report. On a query about a company’s product recalls, the report cites six sources and reads well, but it entirely omits a class action lawsuit that three of the cited sources actually mention in passing. Every subagent’s individual output, checked separately, is accurate. What is the most likely root cause. Option one, the synthesis subagent has a bias toward brevity and is dropping detail. Option two, the coordinator’s initial decomposition scoped the search and analysis subagents toward recall mechanics and safety data specifically. It never scoped them toward legal or financial consequences, so nothing about litigation was ever surfaced as a distinct thread to research. Option three, the document analysis subagent needs a larger context window. Option four, add a fact-checking subagent after synthesis to catch omissions. The answer is option two. Every subagent already returned accurate output for what it was scoped to research. The lawsuit is mentioned in passing in the sources they found. But nothing in either subagent’s task was to identify it as its own subject worth investigating, so it never got escalated into the report as a finding. Option one is attractive, because dropped detail really does happen at synthesis. But the sources here are accurate and cited, not brief or truncated, and that points at what was never asked for rather than what was cut. Option three treats this as an attention problem, and it is a coverage problem, the same distraction from episode one where a bigger window does not fix what a decomposition choice broke. Option four is the most tempting distractor here, because a fact-checking pass sounds like exactly the kind of safeguard a real team would want, and it would catch some classes of error. It still does not fix the actual cause, which is upstream at decomposition. Add a fourth subagent and you have added cost without touching the reason the first two subagents were never pointed at the right question. Three sentences on what the exam will ask. Expect a scenario where every individual subagent succeeded and the defect is still real, and the correct diagnosis names the coordinator’s decomposition rather than any downstream agent. Expect a second item on the mechanics: Task missing from allowedTools, or a subagent behaving as if it remembers something from an earlier turn it was never told. And when parallel execution is on the table, the right answer is usually to spawn the independent calls together in one response. It is not to add a new subagent to catch what the existing ones missed. Next time, the guarantees you cannot leave to a prompt. A refund agent that skips identity verification twelve percent of the time, and why the fix is not a better instruction. This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit brodynetworks.substack.com (https://brodynetworks.substack.com?utm_medium=podcast&utm_campaign=CTA_1)

4 Sept 2026
Episode 1. The Agentic Loop, and Where People Break It
Domain 1, task statements 1.1 and 1.6. An agent that answered correctly on turn four and then ran until the harness killed it at fifty iterations, and the three wrong ways to decide a loop is finished. Then the same question one level up: when to fix your steps in advance, and when to let the model choose them. Independent and unofficial. This series is not affiliated with, sponsored by, or endorsed by Anthropic. Nothing in it is exam content. Every practice question was written for this show against the published exam guide, which is a free public document linked below. Chapters * 0:00 Cold open and disclaimer * 0:24 The agent that would not stop * 1:10 The loop, and what stop_reason is * 2:12 Three things that look like completion signals * 2:26 Wrong signal one: reading the text * 3:12 Wrong signal two: the presence of text * 3:39 Wrong signal three: the iteration cap * 4:27 What stop_reason is not for * 4:46 The stop reasons that also mean keep going * 5:38 Returning tool results into history * 6:28 Model-driven decisions vs a decision tree * 7:38 Task decomposition: fixed vs adaptive * 8:57 A quick test for which shape you have * 9:27 Large code review, and the attention-dilution trap * 10:16 Worked question * 11:12 Why the other three options are there * 12:29 What the exam will ask, and next time Sources * Claude Certified Architect, Foundations Exam Guide, Version 1.0 (section 6, task statements 1.1 and 1.6; section 17 appendix) (https://everpath-course-content.s3-accelerate.amazonaws.com/instructor%2F6nizmqk8tpzpfjvt6qmmav7rh%2Fpublic%2F1783542750%2FClaude+Certified+Architect+%E2%80%93+Foundations+Exam+Guide.pdf) * Tool use overview (`stop_reason`, tool_use content blocks) (https://docs.claude.com/en/docs/agents-and-tools/tool-use/overview) * Handling tool calls (returning tool results into the conversation) (https://platform.claude.com/docs/en/agents-and-tools/tool-use/handle-tool-calls) * Handling stop reasons (the full list, and `pause_turn`) (https://platform.claude.com/docs/en/build-with-claude/handling-stop-reasons) * Claude Agent SDK overview (the agentic loop) (https://docs.claude.com/en/api/agent-sdk/overview) Transcript This is Passing C C A R F, an independent study companion for the Claude Certified Architect, Foundations exam. Sixty items, a hundred and twenty minutes, and a published blueprint that tells you almost exactly what it is going to ask. Independent and unofficial. Not affiliated with or endorsed by Anthropic. No exam content. The agent would not stop. It had answered the customer’s question in the fourth turn, correctly, and then it looked up the same order six more times before the harness killed it at fifty iterations. Nobody had written a bug. Every individual piece worked. The loop just did not know it was done. That is where domain one starts, and it is worth twenty seven percent of this exam, more than any other. Today is task statement one point one, agentic loops, and task statement one point six, task decomposition. Two topics, one underlying idea: something has to decide what happens next, and you need to be deliberate about whether that something is your code or the model. Start with the loop, because it is smaller than people think. You send a request to Claude. You look at what came back. If the model asked to use a tool, you run the tool, you put the result into the conversation, and you send the whole thing again. If it did not ask for a tool, you are finished and you show the user the answer. That is the entire lifecycle. Request, inspect, execute, return, repeat. The inspect step is the one that matters, and it has a specific name. Every response carries a field called stop reason, which tells you why the model stopped generating. When it wants a tool, the value is tool use. When it has finished its turn and is handing the conversation back, the value is end turn. Those two values are your control flow. Continue on tool use. Terminate on end turn. That is the whole rule, and if your loop implements exactly that rule, the failure I opened with cannot happen. So why does it happen so often? Because there are three other things that look like completion signals, and all three are wrong. Let me take them in order of how attractive they are, because the exam will offer you all three as answers. The first is reading the text. The model finishes with something like “let me know if you need anything else”, so you scan the reply for a closing phrase and stop when you find one. This feels reasonable and it is the worst of the three. You have taken a structured, reliable field that the API gives you for free. You have replaced it with natural language pattern matching, against a system whose entire job is generating varied natural language. It will work in testing, because in testing the model says the phrase you wrote your check against. It fails in production the first time the model phrases the ending differently. Or the first time it says it has finished searching, in the middle of a longer job, and your loop exits with two tools still to call. The second is checking whether there is any text in the reply at all. The idea is that text means the model is talking to the user rather than working. That one falls apart immediately, because a model can and routinely does explain what it is about to do and then request a tool in the same response. Text and tool use are not mutually exclusive. Presence of prose tells you nothing about whether the turn is over. The third is the iteration cap, and this is the subtle one, because caps are not bad. Stop after twenty five iterations, the thinking goes, and nothing can run away. A cap is a good thing to have. It is a terrible thing to depend on. The difference is what happens on a normal, healthy request. If your cap is doing its job, it never fires, because end turn already ended the loop. If the cap is your stopping mechanism, every long task ends by being cut off mid work. You have shipped an agent that quietly truncates its own answers whenever a job needs more steps than you guessed. Keep the cap. Make it a circuit breaker, not a design. If it fires, that is an incident to investigate, not a normal Tuesday. One clarification before I move on, because this is where people over apply the rule. Your loop only has to recognise two values. Tool use means keep going. End turn means the loop is done deciding. What you do after that is application logic, sitting outside the loop. But do not turn that into a rule that anything which is not tool use ends the loop. The API has more than two stop reasons, and at least one of the others also means keep going. If you use server side tools, you will meet a value called pause turn, and the documented handling is to send the response back unchanged so the model can finish. A loop that treats every non tool use value as the end will truncate that job silently. That is the same class of bug I opened with, approached from the other side. For the exam, tool use and end turn are the two values in the published objectives, and they are the two your control flow turns on. In production, know that the list is longer. Keeping that boundary clean is why the loop stays four lines long instead of becoming a state machine nobody wants to touch. Now the return step, which is where the second family of bugs lives. When the tool comes back, its result has to go into the conversation history as a message. The whole accumulated conversation then goes back to the model on the next call. This sounds obvious written down. It is easy to get wrong in practice. A natural way to write the code is to run the tool, hold the result in a variable, and ask the model a fresh question about it. Do that and you have thrown away the reasoning that led to the tool call. The model needs its own previous turn, the tool request, and the tool result, sitting together in order. That is what lets it notice that the order lookup returned nothing and try a different identifier, instead of confidently reporting an order that does not exist. Which leads to the part of task statement one point one that I think is the most important idea in the domain. It is a design question, not an implementation one. In a real agentic loop, the model decides which tool to call next, based on what it has learned so far. That is different from a pre configured decision tree, where your code decides the order and the model just fills in the blanks. Both are legitimate architectures. They are not the same architecture, and the exam cares that you know which one you have built. The tell is what happens when reality does not match the plan. In a decision tree, an unexpected result has to be handled by a branch someone thought to write. In a model driven loop, an unexpected result is just context, and the model can respond to it by calling something you did not anticipate it calling. You buy adaptability and you pay in predictability. If you are processing a known form with known fields, the tree is better and cheaper. If you are handling open ended customer requests, the tree will meet a case nobody enumerated within about a day. That trade off is exactly what task statement one point six is about, so let us go there, because it is the same question one level up. Once a job is too big for one pass, how do you break it apart? There are two shapes and the exam wants you to pick between them on purpose. The first is a fixed sequential pipeline, sometimes called prompt chaining. You know the steps in advance, so you write them in advance. Analyze each file, then run a pass that looks across files. Extract the fields, then validate them, then reconcile. The second is dynamic decomposition, where the first step’s findings determine what the second step even is. Map the codebase, see what you found, then decide which areas deserve tests and in what order. Pick the fixed pipeline when the work is predictable and multi aspect. Pick dynamic decomposition when the work is genuinely investigative and you cannot enumerate the subtasks up front without guessing. The failure mode of choosing wrong in one direction is a rigid pipeline that cannot follow a lead. The failure mode in the other direction is an agent that wanders around a well understood problem. It takes six times as long to do something you could have specified in a list. A quick test that settles it most of the time. Try to write down the steps. If you can list them completely and in order before you start, and you would write the same list for the next hundred inputs, you have a fixed pipeline. Build it as one. If step two only becomes knowable after step one runs, you do not have a pipeline. You have an investigation, and forcing it into a fixed shape means guessing at the steps and then living with the guess. There is one specific application of the fixed pipeline that comes up so often it is worth naming on its own. Large code reviews. Hand a model fourteen changed files and ask for a review, and the quality comes back uneven. It is uneven in a particular way. The issues it finds cluster in whichever files it looked at hardest. Split that into a per file pass for local problems, plus a separate pass whose only job is looking across files. You get better coverage from the same model. The reason is attention dilution, and I want to be precise about why that matters, because the exam builds a trap out of it. Attention dilution is not a context window size problem. Making the window bigger does not fix it. The fix is fewer things to attend to per pass. Let me work a question in the shape you will see. A team runs a support agent built on the Agent SDK. Roughly one conversation in twenty never finishes. It keeps calling tools until the runtime stops it at fifty iterations, and the transcripts show the same lookup repeated. Their loop continues while the response contains a tool use block, and it exits when the assistant’s text contains a completion phrase. Which change fixes this. Option one, raise the iteration cap to a hundred so long conversations have room to complete. Option two, terminate the loop when stop reason is end turn, and keep the cap only as a safety limit. Option three, add few shot examples that show the model ending its turn cleanly. Option four, add a hook that blocks a tool call if the same tool was already called with the same arguments. The answer is option two. Now the interesting part, which is why the other three are on the page. Option one is there because it sounds like capacity planning. It is the exam’s favorite kind of wrong answer: it makes the symptom rarer without touching the cause, and it makes every genuine failure take twice as long to surface. Option three is more tempting, because few shot examples are a real technique that really does improve consistency. But it treats a control flow bug as a prompting problem. The model is already signalling end turn correctly. Nobody is reading the signal. No amount of prompting fixes a caller that is not listening. Option four is the most seductive, because deduplicating repeated calls would actually stop the runaway, and hooks are a legitimate enforcement mechanism you will meet again in episode three. It is still wrong, and the reason is the general principle underneath about a third of this exam. Fix the cause at its own layer. A broken termination condition is a loop bug. Patching it downstream with a deduplication hook leaves the broken condition in place, and now you have two mechanisms, one of which is lying about what it is for. Three sentences on what the exam will ask. Expect an item that hands you a loop with a wrong termination condition and asks you to spot it. The distractors will be natural language parsing, checking for text content, and an iteration cap sold as the primary stopping mechanism. Expect a second item that asks you to choose between a fixed sequential pipeline and adaptive decomposition. Answer it by asking whether the subtasks are knowable before you start. And when a bigger context window appears as an option for an inconsistent review, treat it as a distractor. Splitting the work into focused passes is the answer they want. Next time, coordinators and subagents. A research report comes back beautifully written, correctly cited, and missing an entire subject area, every subagent having reported success. We work out why the answer is almost never the subagent that ran last. This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit brodynetworks.substack.com (https://brodynetworks.substack.com?utm_medium=podcast&utm_campaign=CTA_1)

4 Sept 2026
Episode 0. What This Exam Actually Is
Episode zero is the logistics episode: what CCAR-F is, what the sixty items and the 720 cut score actually mean, and how the five domains are weighted. It also covers the part most write-ups skip, which is that registration runs through the Claude Partner Network, and what that means if you do not work at a partner firm. Independent and unofficial. This series is not affiliated with, sponsored by, or endorsed by Anthropic. Nothing in it is exam content. Every practice question was written for this show against the published exam guide, which is a free public document linked below. Chapters * 0:00 Cold open and disclaimer * 0:24 Why this episode is logistics only * 0:36 What the credential is * 1:14 The numbers * 1:40 Why 720 is not 72 percent * 2:42 The five domains and their weights * 3:37 Who the exam is written for * 4:25 Thirty task statements * 4:58 The six-scenario bank * 6:11 Registration, and the Partner Network gate * 7:41 Retakes and the reschedule discrepancy * 8:29 Renewal * 8:59 What is out of scope, so you stop studying it * 10:19 The prep-list trap * 10:51 How this series is built * 11:25 A promise about the questions * 12:05 What the exam will ask, and next time Sources * Claude Certified Architect, Foundations Exam Guide, Version 1.0, effective July 2026 (free public PDF) (https://everpath-course-content.s3-accelerate.amazonaws.com/instructor%2F6nizmqk8tpzpfjvt6qmmav7rh%2Fpublic%2F1783542750%2FClaude+Certified+Architect+%E2%80%93+Foundations+Exam+Guide.pdf) * Certification page (https://anthropic-partners.skilljar.com/claude-certified-architect-foundations-certification) * Official prep course list (https://anthropic-partners.skilljar.com/page/claude-certified-architect-foundations-prep-courses) * Anthropic Partner Academy (https://anthropic-partners.skilljar.com/) * Claude Partner Network application (the ten-employee qualification, checked 2026-09-03) (https://claude.com/form/cpn-partner-application) * Certification announcement (https://claude.com/blog/four-role-based-claude-certifications) * Pearson VUE, Anthropic (the 48-hour reschedule figure) (https://www.pearsonvue.com/us/en/anthropic.html) * Credly badge (https://www.credly.com/org/anthropic/badge/claude-certified-architect-foundations) Transcript This is Passing C C A R F, an independent study companion for the Claude Certified Architect, Foundations exam. Sixty items, a hundred and twenty minutes, and a published blueprint that tells you almost exactly what it is going to ask. Independent and unofficial. Not affiliated with or endorsed by Anthropic. No exam content. Episode zero is logistics. No architecture, no worked question, no code. If you already know what this exam is and how to sit it, skip to episode one. Start with the thing itself. The credential is Claude Certified Architect, Foundations. The exam code is C C A R F, and I will use the code from here on, because the full name does not survive being said forty times. It is a proctored exam delivered by Pearson View. The current exam guide is version one, effective July of twenty twenty six, and it is a free public document. Link in the show notes. Everything I say in this series comes from that guide, from the public Anthropic documentation, and from building the things it describes. Here are the numbers. Sixty items. A hundred and twenty minutes, so two minutes an item if you spend them evenly, which you will not. The pass mark is seven twenty, on a scale that runs from one hundred to one thousand. The fee is a hundred and twenty five US dollars per attempt, and any partner discount shows up at checkout rather than on the page. The credential is good for twelve months from the day it is awarded. Two of those numbers get misread constantly, so let me be specific. Seven twenty out of a thousand is not seventy two percent. It is a scaled score. Anthropic set the bar through a standard setting study. Subject matter experts decided what a just barely qualified candidate should be able to do. The scale exists so scores mean the same thing across exam forms that are not equally difficult. The practical consequence is that you cannot count questions and predict your result. The second misread is the domain breakdown on your score report. You get a pass or fail, a scaled score, and a percentage correct in each of the five content areas. Those per domain percentages are there to tell you where you are weak. They do not decide the outcome. Only the total scaled score does that. So there is no domain you can fail on its own, and there is also no domain you can safely ignore. Now the five domains and what they are worth. Agentic architecture and orchestration is twenty seven percent, the largest single block. Tool design and Model Context Protocol integration is eighteen. Claude Code configuration and workflows is twenty. Prompt engineering and structured output is twenty. Context management and reliability is fifteen. That adds to one hundred, and it is the single most useful table in the guide, because it tells you where your study hours belong. Sit with those weights for a second. Agentic architecture plus Claude Code configuration is forty seven percent of the exam between them. Nearly half the paper is orchestration and tooling judgment rather than API trivia. If your instinct is to go and memorize parameter names, the blueprint is quietly telling you not to. One more thing the guide is explicit about, and it is worth hearing before you book anything. The candidate it describes is a solution architect who builds production applications with Claude. The guide puts that at something like six months or more of hands on time. Across the API, the Agent SDK, Claude Code and Model Context Protocol. The guide names no formal prerequisites. But the word foundations in the title refers to the scope of the material, not to the difficulty of the questions, and that reading is mine rather than the guide’s. This is a foundations exam in the sense that it covers the base layer of the whole stack. It will still ask you what to do when a production system is misbehaving, in a way you have to reason about. Underneath the domains sit thirty task statements. Seven in domain one, five in domain two, six each in domains three, four and five. Each one lists what you are expected to know and what you are expected to be able to do. The guide says items are written against those objectives, and you should read that sentence literally. The task statements are the specification for the question bank. If a topic is not in a task statement, it is not on the exam, and if it is in one, it is fair game. The other structural thing to understand is the scenario bank. This is not sixty standalone questions. The exam is built from six production scenarios, and on the day you get four of them, picked at random. A customer support resolution agent with backend tools and an escalation path. Code generation with Claude Code. A multi agent research system with a coordinator and specialist subagents. A developer productivity agent working over an unfamiliar codebase. Claude Code wired into a continuous integration pipeline. And a structured extraction system pulling fields out of unstructured documents. That structure changes how you should study, and this is the part I would underline. You cannot revise for four of six. You have to be ready for all six, because you do not know which four you get. But it also means the questions are contextual rather than abstract. You will not be asked to define a hook. You will be handed a production system behaving badly, and the answer will often turn on whether the requirement is deterministic or probabilistic. Study the judgment, not the vocabulary. Now, registration, which is where this gets awkward. The chain has three links. You register through the Anthropic Partner Academy, you complete checkout, then you create a Pearson View account and schedule the session, online proctored or at a test center. That is the published path and it works fine, right up until the first link. The Partner Academy is not the public Anthropic academy. It requires additional validation at login, and Anthropic describes these certifications as available to members of the Claude Partner Network. So the real prerequisite for this exam is not six months of experience. It is being inside a partner firm. And when I checked the partner application form, one of the listed qualifications was being a registered business with ten or more employees. I want to be careful here, because that is a live web page and not a term in the exam guide, and pages change. Check it yourself before you plan around it. But as things stand, an independent consultant cannot simply pay a hundred and twenty five dollars and book a seat. If you work at a partner firm, you already have the access and this is a non issue. If you do not, the credential is currently out of reach, and the honest thing for me to say is that I am in the second group. I am not going to sit this exam. Which means, usefully, that nothing in this series can ever be contaminated by having seen a real question. If it goes badly on the day, there is a retake ladder. Fourteen days after a first failure, thirty after a second, ninety after a third, and a cap of four attempts in a rolling twelve month period. The fee applies every time. Those limits are per exam, so failing this one does not lock you out of a different Anthropic certification. One discrepancy worth carrying in your head. The exam guide says you can cancel or reschedule up to twenty four hours before your appointment, and that changes inside that window cost you the fee. The Pearson View page for Anthropic says forty eight hours for test center appointments. Two official sources, two numbers. Assume forty eight and you are never wrong. No shows forfeit the fee and have to register again. Renewal is better news than most certification programs manage. The credential expires after twelve months, but if you renew on time you review what has changed and take a free assessment on the Partner Academy, not proctored, no fee. Let it lapse and you retake the whole thing at full price. And if the exam content changes substantially, Anthropic reserves the right to make holders retake the full exam anyway rather than take the short assessment. Now the section that saves you the most hours, which is what the guide says will not be on the exam. Fine tuning and training custom models. API authentication, billing and account management. Deep implementation detail in any particular language or framework, beyond what tool and schema configuration needs. Deploying and hosting Model Context Protocol servers, including anything to do with infrastructure or containers. Claude’s internal architecture and training. Constitutional A I, reinforcement learning from human feedback, and safety training methods. Embeddings and vector databases. Computer use. Vision and image analysis. Streaming and server sent events. Rate limits, quotas and pricing math. Authentication protocols and key rotation. Cloud provider specifics for the big three. Benchmarking and model comparison. Prompt caching beyond knowing that it exists. And token counting and tokenization internals. Read that list twice. It is long, it is specific, and a great deal of the general purpose Claude study material on the internet is about things on it. There is a trap attached. Anthropic publishes a recommended prep list of seven free courses, and two of them are about running Claude on Google Cloud and on Amazon Bedrock. Cloud provider specifics are on the out of scope list. I am not telling you the official prep list is wrong, and those courses are genuinely useful if that is your deployment target. I am telling you that finishing them will not move your score, so if your study time is tight, they go last. Which brings me to how this series is built. Thirteen episodes. This one, then eleven that walk the blueprint, then a final episode on exam day and on how the questions are actually constructed. Runtime is allocated by domain weight rather than evenly, so the twenty seven percent domain gets three episodes and the fifteen percent domain gets two. Every episode after this one opens on something failing in production. Then the concepts, then the angle the exam is likely to come at them from. Then a worked question and a short recap. One promise about the questions. Every practice question in this series was written by me, against the published objectives. None of them came from anyone who has sat the exam, and none of them ever will. The guide contains a set of sample questions that Anthropic published openly, and those are the only pre existing items this show will ever discuss, in the final episode. If you find a site advertising real exam questions and daily feedback from test takers, close the tab. Those are braindumps, using them can cost you the credential, and quoting them here would make this show part of that ecosystem. So, what the exam will ask about episode zero. Nothing. There is no item on the paper about fees, scaled scores or reschedule windows. The only reason to know any of it is so none of it surprises you on the day. The material starts in episode one. Next time, the agentic loop, and the three ways people break it. We start with an agent that would not stop running. We work out why reading the text of a reply is the wrong way to know it has finished. Then we get somewhere useful about when to fix the steps in advance and when to let the model choose them. One piece of housekeeping before I go. This series is independent. It is not affiliated with, sponsored by, or endorsed by Anthropic. Nothing here is exam content. Every question you hear was written by us against the published exam guide, which is a free public document, and I will link it in the notes. If you want the authoritative source, that guide is it, and this show is a companion to it, not a replacement for it. This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit brodynetworks.substack.com (https://brodynetworks.substack.com?utm_medium=podcast&utm_campaign=CTA_1)
Contact Passing CCAR-F
- Guest appearances
- Does not typically book guests
Based on episode analysis; this does not confirm that the show is currently accepting guests.
Host of Passing CCAR-F?
Claim your podcast to manage its listing and keep your show details accurate.
Pod Engine is an independent podcast discovery and analytics service and is not affiliated with or endorsed by this podcast. Artwork and show content belong to their owners. Full legal notice.
Explore this show
Podcast research with Pod Engine