{"id":2684,"date":"2026-07-21T12:54:00","date_gmt":"2026-07-21T12:54:00","guid":{"rendered":"https:\/\/futurenews24.com\/index.php\/2026\/07\/21\/atom-everything-31\/"},"modified":"2026-07-22T06:59:40","modified_gmt":"2026-07-22T06:59:40","slug":"atom-everything-31","status":"publish","type":"post","link":"https:\/\/futurenews24.com\/index.php\/2026\/07\/21\/atom-everything-31\/","title":{"rendered":"A Fireplace Chat with Cat and Thariq from the Claude Code crew"},"content":{"rendered":"<p><br \/>\n<\/p>\n<div data-permalink-context=\"\/2026\/Jul\/21\/cat-and-thariq\/\">\n<h2>A Fireplace Chat with Cat and Thariq from the Claude Code crew<\/h2>\n<p class=\"mobile-date\">twenty first July 2026<\/p>\n<p>Earlier this month I hosted a hearth chat session on the AI Engineer World\u2019s Truthful with Cat Wu and Thariq Shihipar from Anthropic\u2019s Claude Code crew. We talked about Claude Code, Claude Tag, Fable, coding agent safety, evals, software design, and the way Anthropic use these instruments themselves.<\/p>\n<p>The complete video of the session is now out there on YouTube. Under is an edited copy of the transcript, with further hyperlinks and my very own bolded highlights.<\/p>\n<p>A number of top-level notes in the event you don\u2019t wish to watch the video or wade by way of the entire transcript:<\/p>\n<p>Claude Tag (Claude\u2019s new collaborative Slack integration) now lands 65% of the product engineering PRs for the Claude Code crew.<br \/>\nClaude Code ships options to Anthropic workers first, and solely ships the options that reveal consumer retention with that cohort<\/p>\n<p>Crucial modifications to Claude Code are nonetheless reviewed manually, however the crew more and more depends on automated code evaluate for the \u201couter layers\u201d of the product.<br \/>\nIncluding examples to a system immediate is now not finest observe for fashions like Fable 5 and even Opus 4.8. The Claude Code system immediate just lately shriveled by 80%.<br \/>\nLikewise, lists of &#8220;don\u2019t do X and don\u2019t do Y&#8221; can scale back the standard of outcomes from the newest fashions.<\/p>\n<p>Dogfooding inside Anthropic is named &#8220;ant fooding&#8221;.<br \/>\nAnthropic actually consider of their auto mode, and see that as an enabling expertise for Claude Tag.<br \/>\nThariq advises offsetting coding-agent-induced Deep Blue by &#8220;being extra bold&#8221; with the work you tackle.<br \/>\nFable is competent at modifying video, and Thariq used it to edit its personal launch video.<br \/>\nAnthropic\u2019s tradition of working (internally) in public is vital to their success, as demonstrated by the way in which they use Claude Tag of their public Slack Channels.<\/p>\n<h4 id=\"how-has-what-you-do-day-to-day-changed-in-the-past-year-\">How has what you do day-to-day modified up to now 12 months?<\/h4>\n<p>1:05<\/p>\n<blockquote>\n<p>Simon: Claude Code got here out in February of final 12 months \u2014 it\u2019s beneath a 12 months and a half previous, and it was initially only a bullet level on the Claude Sonnet 3.7 launch. How has what you do on a day-to-day foundation modified up to now 12 months, now that we&#8217;ve these coding brokers that truly work for us?<\/p>\n<p>Cat: I bear in mind once we first got here out with Claude Code and Sonnet 3.7, you&#8217;d give it a job and you would need to intently monitor each single little factor it tried to do. I might learn each permission immediate extraordinarily fastidiously. I might ceaselessly say no \u2014 no, no, no, did you test this file? Did you test that file? And now it\u2019s been unimaginable with each mannequin technology. I really feel like we\u2019ve all gotten an opportunity to take a step again and delegate much more of the menial implementation to Claude. It\u2019s freed up numerous our time to consider extra inventive work, like: what&#8217;s the proper expertise that we needs to be offering to our customers, now that we all know Claude Code can implement numerous it? And now with Fable it\u2019s a completely completely different step change enchancment. We see for lots of our use circumstances you can really one-shot a ton of options with Fable now.<\/p>\n<p>Thariq: I bear in mind the primary textual content I bought about Claude Code. One in every of my finest pals was like, \u201cYou could go strive Claude Code.\u201d It was about when Opus 4 got here out, and I attempted it and I used to be like, \u201cOh, shit. I must work at Anthropic now.\u201d And that was Opus 4 \u2014 nice mannequin, however you had been studying permission prompts. It\u2019s sort of loopy how a lot amnesia we&#8217;ve, the place I\u2019m like, oh, auto mode has at all times been right here, proper? I don\u2019t even bear in mind urgent sure and permit. For me, the massive factor I\u2019m attempting to push myself on is that we&#8217;ve to do greater high quality work than we\u2019ve ever performed earlier than. The outputs are extremely prime quality. I\u2019ve been utilizing it to edit movies a bunch, and I\u2019m like, okay, it has to satisfy the very exacting calls for of our model crew in a few hours or we simply can\u2019t do it. That\u2019s how I\u2019m attempting to shift with Fable: the most effective work we\u2019ve ever performed, sooner than we\u2019ve ever performed it earlier than.<\/p>\n<\/blockquote>\n<h4 id=\"what-piece-of-conventional-software-engineering-no-longer-holds-\">What piece of typical software program engineering now not holds?<\/h4>\n<p>3:39<\/p>\n<blockquote>\n<p>Simon: What\u2019s a chunk of typical software program engineering that was true a 12 months in the past that you just don\u2019t assume holds anymore on this new world?<\/p>\n<p>Cat: One of many largest shifts we\u2019re seeing within the eng ability set: two years in the past it was fairly typical for a product supervisor to go discuss to a bunch of shoppers, align over the course of six months with cross-functional groups on some PRD, and write a radical spec on precisely how we\u2019ll implement this earlier than the primary line of code will get written. Now issues are utterly turned the alternative approach. For lots of engineers, the push I might give to people within the room is to develop extra of your corporation sense and product sense on what it&#8217;s we must always construct, as a result of the timeline between having an thought and constructing it&#8217;s so a lot shorter \u2014 it\u2019s down from six to 12 months to possibly even per week. Which means all of us must have higher style on what&#8217;s price constructing, what is going to really inflect the companies we\u2019re engaged on. So it\u2019s a rise in worth on product style and enterprise sense, and a bit decrease on execution in most product domains. After all, for infra there\u2019s nonetheless a really heavy emphasis on ensuring all the small print are proper.<\/p>\n<p>Thariq: For me, it\u2019s that rewrites at the moment are good.<\/p>\n<p>Simon: The worst factor you would do is now really wonderful!<\/p>\n<p>Thariq: Precisely. All of the Legendary Man-Month stuff \u2014 by no means rewrite \u2014 I\u2019m pro-rewriting now. If in case you have  check suite \u2014 and I feel the rewrite really forces you to ensure you have  check suite \u2014 however I feel what folks undercount is {that a} codebase is a spec, and possibly it\u2019s the one copy of the spec that you&#8217;ve got, as a result of nobody is aware of each branching a part of the codebase. You&#8217;ll be able to take this as an artifact and distill it or create different variations of it. We rewrote Bun in Rust and it really works nice \u2014 it\u2019s reside for me proper now.<\/p>\n<p>Simon: You\u2019re not transport Claude Code on Bun-in-Rust but, proper?<\/p>\n<p>Thariq: Internally we&#8217;ve.<\/p>\n<\/blockquote>\n<p>(Really it appears like Anthropic began transport Claude Code on Bun-in-Rust to everybody on June seventeenth.)<\/p>\n<h4 id=\"what-kind-of-things-are-non-engineers-doing-with-claude-tag-\">What sort of issues are non-engineers doing with Claude Tag?<\/h4>\n<p>6:36<\/p>\n<blockquote>\n<p>Simon: The opposite huge launch just lately was Claude Tag \u2014 that\u2019s what, per week previous now, a minimum of for the remainder of us. I perceive it\u2019s getting used at Anthropic by non-engineers a terrific deal. What sort of issues are non-engineers doing with Claude Tag?<\/p>\n<p>Cat: Claude Tag is a Claude that lives in your crew\u2019s collaboration instruments. We launched it final week inside Slack. The factor that\u2019s completely different about Claude Tag is it\u2019s multiplayer by default. When you add Claude Tag to a Slack channel, you possibly can chime in, your teammates can chime in, and you may collaborate collectively on the PR. The opposite huge distinction is that it\u2019s proactive as an alternative of reactive. You&#8217;ll be able to inform Claude Tag, \u201cHey, monitor each bug report on this channel, put up a PR to repair it, and tag the engineer who final touched this a part of the codebase,\u201d and it\u2019ll do it for the lifetime of the channel with out you having to manually tag it in. And the third huge shift is that we\u2019ve added crew reminiscence into this. In the event you inform Claude Tag your preferences within the channel, it\u2019ll bear in mind them for each future publish. In the event you at all times need it to debug outages however you don\u2019t need it to debug warnings, simply inform it that in pure language within the channel and it\u2019ll bear in mind it for you and everybody else in your crew.<\/p>\n<p>Internally, we see Claude Tag because the evolution of Claude Code. We see this as a big shift in how we work internally. Claude Tag at the moment lands 65% of our product eng PRs.<\/p>\n<p>Simon: For all of Anthropic, or simply for Claude Code?<\/p>\n<p>Cat: That is only for our product engineering crew \u2014 our inner model of Claude Tag lands 65% of our product PRs proper now. And it is a big shift; that is greater than 50% of our PRs. The best way we see folks cut up work between Claude Code and Claude Tag is: Claude Code remains to be the most effective place to your most complicated duties, if you\u2019re interactively iterating with the agent. However Claude Tag is nice for having it work proactively in your behalf, so that you now not must manually kick off Claude Code for all of the bug stories that come up for options you\u2019re engaged on.<\/p>\n<p>Thariq: And for non-coding circumstances: for instance, earlier than this discuss we requested Claude Tag, \u201cHey, when is Fable releasing?\u201d We needed to ensure we\u2019d line it up with the announcement. Claude Tag would search our Slack and take a look at who\u2019s been saying what. As a search engine to your firm, it\u2019s actually priceless. It has all of the context to your product, so you possibly can ask it metrics-related questions \u2014 usually if you\u2019re making selections you need them knowledgeable by what the metrics say, so that you hook it as much as your occasion retailer. I\u2019ve seen our advertising and marketing crew do issues like, &#8220;Hey, inform me about this characteristic.&#8221; They\u2019re not programmers, however Claude is a programmer \u2014 it could actually clone the codebase and say, &#8220;That is the characteristic, that is what it appears like, it is a recording of me utilizing the characteristic.&#8221; It allows a complete huge number of issues, and I feel we\u2019re nonetheless early in figuring that out.<\/p>\n<\/blockquote>\n<h4 id=\"claude-tag-as-the-team-collaborative-layer\">Claude Tag because the crew collaborative layer<\/h4>\n<p>10:06<\/p>\n<blockquote>\n<p>Simon: One of many issues I\u2019ve had with coding brokers is that I get learn how to use them as a person, however I\u2019m probably not clear on learn how to use them in a crew setting. It feels like Claude Tag is your present reply to that crew collaborative layer for these items.<\/p>\n<p>Cat: Precisely. And a big proportion of our classes are literally multiplayer proper now. Perhaps I say, \u201cHey, I feel we must always implement this new characteristic in Cowork,\u201d and I\u2019ll tag in Claude Tag to do a primary cross at it. Then I\u2019ll inform Claude Tag, \u201cShare a recording of your remaining implementation,\u201d and I\u2019ll tag in design to have a look. They\u2019ll nudge it, then cross it on to eng to take it to the end line and get it out to prod. It\u2019s been this very fluid expertise. We\u2019re nonetheless attempting to iron out what the social dynamics are for steering the identical session, however we\u2019ve discovered that individuals simply observe how others use it and comply with these social norms \u2014 it\u2019s been fairly intuitive for us to combine Claude Tag into our groups.<\/p>\n<p>Thariq: It\u2019s nice for educating folks, and in addition for lowering slop, as a result of the truth that everyone seems to be seeing you employ Claude collectively type of ranges up how you employ Claude as properly.<\/p>\n<\/blockquote>\n<p>This jogged my memory of how Midjourney solved the problem of educating folks superior picture prompting by implementing prompting in public of their Discord channels.<\/p>\n<h4 id=\"how-do-you-decide-which-features-are-worth-building-when-building-is-so-much-cheaper-\">How do you determine which options are price constructing when constructing is a lot cheaper?<\/h4>\n<p>11:41<\/p>\n<p>One thing I\u2019ve discovered actually arduous myself is understanding when a characteristic is price transport now that the price of really constructing options has dropped a lot.<\/p>\n<blockquote>\n<p>Simon: How do you take care of the toughest downside in all of engineering \u2014 prioritization? How do you determine which options are price constructing and transport when constructing a characteristic is a lot extra cheap now?<\/p>\n<p>Cat: That is the arduous factor. There are a number of methods we strategy it. One is we dogfood our merchandise each single day. At any time when there\u2019s one thing we wish to have the ability to do in our merchandise that we\u2019re not capable of, as an alternative of discovering a unique answer we repair our product so it could actually assist that case. Now we have a really heavy dogfooding tradition internally. Earlier than we share our merchandise with everybody on this planet, we share them with everybody inside Anthropic, and with some early clients who give us very sincere suggestions about it \u2014 the extra brutal the higher \u2014 and we iterate till folks find it irresistible. Now we have an inner bar for the variety of energetic customers and the quantity of retention a characteristic has to have earlier than we share it with the world. As a result of this bar could be very clear, each engineer is aware of what they\u2019re attempting to hit. I feel this additionally ranges up our polish, as a result of if the characteristic isn\u2019t polished, folks will churn \u2014 after which we shouldn\u2019t ship that characteristic.<\/p>\n<\/blockquote>\n<p>Utilizing inner user-retention to determine if a characteristic ought to ship makes a complete lot of sense to me.<\/p>\n<h4 id=\"do-you-have-an-example-of-a-feature-which-surprised-you-\">Do you might have an instance of a characteristic which shocked you?<\/h4>\n<p>12:54<\/p>\n<blockquote>\n<p>Simon: Do you might have an instance of a characteristic which shocked you? You rolled it out and the engagement was off the charts \u2014 one thing unlikely to be shipped that changed into an actual product factor.<\/p>\n<p>Cat: I do have one. A number of people on our crew love distant management. Distant management enables you to use your cell gadget, or Claude within the internet browser, to connect with an area Claude Code session operating in your CLI. I by no means have this want, as a result of I simply kick off the duty straight on cell and it runs in a cloud session with out utilizing my native setting \u2014 I feel as a result of I\u2019m doing very straightforward coding duties. It was one thing I didn\u2019t completely perceive; I used to be like, hey, folks ought to simply arrange distant dev environments. However in observe, as soon as we rolled out distant management, so many individuals I discuss to informed me that what they do each night time is plug their laptop computer into an influence charger, open a bunch of distant management classes, lock the display, after which use their cell phone from their sofa to manage Claude Code. So this has develop into a circulate we\u2019re now leaning into that I didn\u2019t initially get \u2014 however now I do.<\/p>\n<\/blockquote>\n<h4 id=\"does-a-human-review-every-line-of-production-code-in-claude-code-\">Does a human evaluate each line of manufacturing code in Claude Code?<\/h4>\n<p>14:20<\/p>\n<p>One of many over-arching themes of the convention was evaluate: how a lot consideration to folks spend to reviewing code written for them by coding brokers. I used to be very eager to listen to the Claude Code crew\u2019s tackle this!<\/p>\n<blockquote>\n<p>Simon: How does code evaluate work? Does a human being evaluate each line of manufacturing code that makes it into Claude Code? And if not, what are you doing \u2014 how do you retain the standard up?<\/p>\n<p>Thariq: It varies on the duty quite a bit. For vital areas we&#8217;ve code house owners. The system immediate is an instance the place we&#8217;ve a code proprietor \u2014 you really want to get their approval.<\/p>\n<p>Simon: So the code proprietor is straight liable for the standard of that space of the code.<\/p>\n<p>Thariq: That\u2019s proper.<\/p>\n<p>Cat: And they should approve any PR that touches it.<\/p>\n<p>Thariq: Now we have our code evaluate GitHub bot evaluate every little thing \u2014 that goes on each PR, and infrequently it\u2019s doing the majority of the evaluate. One thing I\u2019ve seen on the crew is that for extra complicated PRs you would possibly make an artifact to elucidate the PR in order that different folks can then evaluate. And we make investments quite a bit into verification, CI\/CD, issues like that, to ensure that any time something fails we&#8217;ve a check. Now we have a very strong setting the place Claude can management Claude Code and check it. So there\u2019s a multi-pronged strategy to code evaluate.<\/p>\n<p>Cat: Generally, we are attempting to maneuver to a world the place people don\u2019t must be within the loop. For probably the most vital modifications to the core of Claude Code, and the cores of different merchandise, there&#8217;s at all times a code proprietor and so they do manually evaluate all of the modifications. However more and more, for the modifications on the outer layers, we even have Claude code evaluate totally evaluate these. That sounds fairly scary, however we\u2019ve had a six-plus-month-long course of to get right here, and there are child steps that you just take to construct up belief with code evaluate. At first we had human evaluate for every little thing, after which more and more we might say, okay, for code modifications that contact these recordsdata, code evaluate is catching 100% of the problems there \u2014 so we really don\u2019t want a human manually reviewing these. And when we&#8217;ve incident evaluate, we take a look at the PRs that triggered the incident and say, okay, how will we replace code evaluate to catch that? \u2014 and we take these PRs and add them to an eval set to ensure our future modifications to code evaluate by no means regress that metric. Eradicating people from the code evaluate loop is an enormous step ahead. It might probably sound scary, and it\u2019s not one thing you are able to do in a single day, however it&#8217;s one thing you are able to do by way of many months of funding within the infrastructure to provide the confidence that code evaluate is catching every little thing you care about.<\/p>\n<\/blockquote>\n<p>So the important thing appears to be continually iterating on the automated evaluate programs themselves, with a view to construct belief in them over time.<\/p>\n<h4 id=\"how-does-a-new-model-affect-your-intuition-for-what-it-can-and-can-t-do-\">How does a brand new mannequin have an effect on your instinct for what it could actually and might\u2019t do?<\/h4>\n<p>17:20<\/p>\n<p>We bought deep into evals\u2014one other sizzling matter all through the broader convention.<\/p>\n<blockquote>\n<p>Simon: I do know that Opus 4.8, if I ask it to construct me a JSON endpoint that runs a SQL question and outputs JSON, is simply going to get it proper \u2014 that\u2019s not one thing I&#8217;ve to evaluate intently. However then a brand new mannequin comes alongside and I don\u2019t know learn how to construct belief in Fable shortly, that it\u2019s not going to mess issues up that Opus didn\u2019t. How does the brand new mannequin have an effect on your instinct for what it could actually do and what it could actually\u2019t do?<\/p>\n<p>Cat: The primary purpose we\u2019re increase this eval base over time is in order that new fashions generally is a drop-in alternative. When we&#8217;ve a brand new mannequin, we run the entire eval set and ensure that, for instance, Fable is strictly higher than Opus 4.8 \u2014 and that provides us the boldness to drop it in.<\/p>\n<p>Simon: Are these mannequin evals for Anthropic as a complete, or Claude Code team-specific?<\/p>\n<p>Cat: Now we have each. Now we have evals on our crew, and we run code evaluate throughout each repo inside Anthropic, so we&#8217;ve evals for that. And for issues like auto mode, we not solely have evals throughout each consumer inside Anthropic \u2014 we\u2019ve additionally commissioned a number of exterior testers to purple crew it, to create environments with immediate injections and malicious inputs, and ensure that auto mode doesn\u2019t let any of these cross.<\/p>\n<\/blockquote>\n<h4 id=\"how-do-you-build-confidence-that-a-system-prompt-tweak-results-in-better-output-\">How do you construct confidence {that a} system immediate tweak leads to higher output?<\/h4>\n<p>18:41<\/p>\n<blockquote>\n<p>Simon: I wish to know if the system immediate enchancment I made really improved the product \u2014 that\u2019s probably the most primary type of product-specific eval, and I nonetheless don\u2019t have a terrific really feel for the way to try this. Is that one thing you\u2019re doing such that you&#8217;ve got full confidence {that a} tweak you\u2019ve made to the system immediate leads to higher output?<\/p>\n<p>Cat: We don\u2019t have full confidence, however we do quite a bit to ensure that we don\u2019t regress efficiency. The start line is a collection of exterior evals that we belief, and we complement that with an excellent bigger suite of inner evals that we belief. To begin, we primarily optimize for functionality: given an entire definition of a job and the complete codebase, does Claude make the fitting selections, totally repair the bugs, and cross all of the exams? That\u2019s the start line and the factor we optimize for, as a result of it\u2019s most straight what customers need. However there are numerous behaviors that affect how customers really feel after they work with Claude Code. For instance, folks actually don\u2019t prefer it when Claude Code says it\u2019s time to fall asleep. Or folks actually don\u2019t prefer it when it says, \u201cHey, I completed two out of 5 elements \u2014 would you like me to proceed?\u201d Sure, please proceed. So we\u2019re increase a set of behavioral evals to catch these. And as we get consumer suggestions \u2014 please be loud with us about your consumer suggestions \u2014 we rank the precedence points and go down one after the other and construct evals for every of them. It\u2019s not 100% protection, however it&#8217;s a precedence for us to extend the protection.<\/p>\n<\/blockquote>\n<h4 id=\"how-much-interaction-is-there-between-the-claude-code-team-and-the-model-training-teams-\">How a lot interplay is there between the Claude Code crew and the mannequin coaching groups?<\/h4>\n<p>20:21<\/p>\n<blockquote>\n<p>Simon: How a lot interplay is there between the Claude Code crew and the groups at Anthropic who&#8217;re coaching the fashions within the first place? Is that fairly a detailed collaboration?<\/p>\n<p>Cat: Throughout Anthropic, all of us work fairly intently collectively. We meet usually to speak about what we count on the following technology of fashions to have the ability to do. Our analysis crew has additionally been wonderful about displaying this publicly \u2014 we frequently discuss in our weblog posts about how we\u2019re concentrating on ever-increasing longer-horizon work, and the way we practice Claude itself to be sincere, innocent, and useful. We additionally put numerous effort into ensuring it\u2019s aligned together with your intent, even when your intent is expressed in a fuzzy approach. After all, strive your finest to be particular about what you need, so Claude has all of the context \u2014 however even if you\u2019re not particular, we train Claude to make good assumptions. It\u2019s been a productive partnership.<\/p>\n<\/blockquote>\n<h4 id=\"the-system-prompt-has-been-reduced-by-80-what-have-you-been-able-to-drop-\">The system immediate has been lowered by 80% \u2014 what have you ever been capable of drop?<\/h4>\n<p>21:24<\/p>\n<p>So many helpful prompting suggestions on this part!<\/p>\n<blockquote>\n<p>Simon: Thariq, you talked about this morning that the system immediate for Claude Code has been lowered by 80% due to Claude Fable. Are you able to go into a bit of extra element? What sort of issues have you ever been capable of drop?<\/p>\n<p>Thariq: It wasn\u2019t simply Fable \u2014 it was Opus 4.8 as properly, and going ahead, future fashions. Now we have completely different system prompts for various fashions now. One of many patterns we noticed is that we had been over-constraining Claude. The preliminary, possibly Opus 4-ish fashions needed numerous examples, and eradicating examples was extraordinarily useful, as a result of it was simply extra inventive than the examples we gave it.<\/p>\n<p>Simon: That\u2019s actually fascinating, as a result of one of many high prompting suggestions I give folks is: give it examples. If that\u2019s now not true, that sort of breaks my prompting mannequin a bit of bit.<\/p>\n<p>Thariq: Similar right here \u2014 I used to be shocked to listen to that. I feel now it\u2019s extra concerning the form of what you give it \u2014 the instruments you give to Claude, your system immediate, issues like that. The opposite factor we did is attempt to give it extra context and fewer \u201cdon&#8217;t do that\u201d directions, as a result of that\u2019s a really robust impulse for Claude, and particularly if it conflicts with consumer directions afterward, that may be extraordinarily complicated to Claude \u2014 \u201cI\u2019ve bought this ability that claims this and the system immediate says this.\u201d So we attempt to have fewer arduous constraints, extra context, and fewer directions general. It\u2019s positively a science \u2014 it took a bunch of evals to construct.<\/p>\n<p>Cat: Generally, if you\u2019re prompting these fashions, it is best to at all times assume: are there edge circumstances to the instruction that I\u2019m giving it? After we went again and reviewed all of the directions within the Claude Code system immediate, we discovered a number of circumstances the place sure, this assertion is 90% true, however there\u2019s an actual 10% of circumstances the place it\u2019s not true. We didn\u2019t wish to constrain the mannequin, or confuse it into considering it ought to at all times do that. One good instance is verification. Everybody right here desires Claude to confirm its work, and we had some directions within the immediate that stated: in the event you make a front-end change, at all times confirm. However there\u2019s a restrict to it. If it\u2019s altering copy from one string to a different string, and the consumer says \u201csimply make a fast repair and replace the check,\u201d possibly you don\u2019t wish to confirm. So we\u2019ve adjusted our wording from \u201cat all times confirm, confirm, confirm\u201d to one thing like: more often than not if you\u2019re doing front-end work you possibly can\u2019t totally perceive the expertise by hitting the backend endpoints, so if you make bigger modifications to the consumer expertise, please run the app domestically. And actually, that instruction in all probability isn\u2019t even good both, as a result of what&#8217;s a big change? Perhaps it ought to check small modifications too. Generally, everytime you give a immediate to the mannequin, it is best to take into consideration the methods during which it may very well be misinterpreted by a well-intentioned human, with a view to higher perceive how the mannequin would possibly interpret it \u2014 and soften the immediate in order that it\u2019s really 100% correct, since you\u2019re giving this immediate to the mannequin 100% of the time.<\/p>\n<p>Simon: What\u2019s fascinating about that&#8217;s you\u2019re counting on the mannequin\u2019s judgment \u2014 and that\u2019s bought to be an Opus\/Fable-level factor. Fashions a 12 months in the past didn&#8217;t have the extent of judgment essential to determine whether or not they had been going to check a change or not. However that does break down in the event you\u2019re constructing for a variety of fashions and attempting to run the cheaper fashions for cheaper duties.<\/p>\n<p>Cat: We even have a unique system immediate per mannequin now, for this very purpose. It\u2019s solely our most frontier fashions which have this 80% token lower \u2014 the older fashions nonetheless have the complete system immediate.<\/p>\n<p>Simon: Do you assume Fable and Opus are good sufficient to immediate Haiku with extra particulars, as a result of they perceive that Haiku has much less judgment, much less style?<\/p>\n<p>Cat: We haven\u2019t been capable of eval it \u2014 we don\u2019t have any arduous information to indicate it.<\/p>\n<p>Thariq: There\u2019s a troublesome factor with smaller fashions typically, as a result of typically the bigger fashions may be extra token-efficient on a tough downside than the smaller fashions. So there\u2019s a little bit of instinct to construct there \u2014 typically you actually simply need frontier intelligence nearly on a regular basis. The Pareto curve shifts, and it\u2019s arduous to seek out.<\/p>\n<p>Simon: A 12 months in the past I didn&#8217;t belief a mannequin to put in writing a immediate. At the moment the nice fashions are superb at prompting \u2014 numerous my prompts are written by fashions, which feels absurd however works very well. What helped me come to phrases with that was fascinated by subagents, that are fully a couple of Claude mannequin organising a immediate for one more Claude mannequin.<\/p>\n<p>Thariq: Workflows are literally a very good instance of this, as a result of it\u2019s Claude not simply prompting a single subagent, however prompting the orchestration of many subagents, and every one in every of them will get a really detailed immediate. It\u2019s nearly a degree above simply spawning a subagent. I\u2019ve additionally been utilizing it on my private machine, giving it the Gemini API and saying: right here, generate photographs. It\u2019s approach much less lazy than I&#8217;m at prompting a picture mannequin. It\u2019s simply Claude prompting Claude all the way in which down.<\/p>\n<p>Cat: I feel Claude additionally wrote the immediate for the workflow software.<\/p>\n<p>Simon: I\u2019ve learn that immediate \u2014 it\u2019s  immediate. That\u2019s really a frustration I&#8217;ve with Anthropic typically: you publish the prompts for Claude Chat, however you don\u2019t embrace the software prompts and the Claude Code prompts. I nonetheless need to run a proxy to intercept them. I might find it irresistible if the Claude Code prompts had been intentionally revealed \u2014 they\u2019re the documentation. They\u2019re how  what the software can do and the way it works.<\/p>\n<p>Cat: I\u2019ll write down that characteristic request. I\u2019ll have Claude Tag do it.<\/p>\n<\/blockquote>\n<p>Fascinating to notice that OpenAI\u2019s prompting finest practices for GPT-5.6 contains comparable recommendation for his or her newest fashions:<\/p>\n<blockquote>\n<p>Favor leaner prompts<\/p>\n<p>Eradicating repeated directions and examples and simplifying software descriptions can enhance job efficiency and token effectivity. In a pattern of inner coding-agent eval runs, configurations with leaner system prompts improved analysis scores by roughly 10\u201315% whereas lowering whole tokens by 41\u201366% and price by 33\u201367%.<\/p>\n<\/blockquote>\n<h4 id=\"what-s-your-bar-for-introducing-a-new-tool-\">What\u2019s your bar for introducing a brand new software?<\/h4>\n<p>28:06<\/p>\n<blockquote>\n<p>Simon: Claude Code is principally an enormous bag of instruments. What\u2019s your bar for introducing a brand new software? How do you determine when it\u2019s price doing that further engineering at that degree?<\/p>\n<p>Cat: Do you wish to take it? You launched among the best instruments we&#8217;ve.<\/p>\n<p>Thariq: My profession peaked after I launched the ask consumer query software. It\u2019s actually arduous. Particularly for some instruments \u2014 ask consumer query is Claude\u2019s software to ask you \u2014 so it\u2019s arduous to eval, and typically it\u2019s extra of a consumer desire factor. Again then we had fewer evals, so it was very dogfooding based mostly \u2014 or \u201cant fooding,\u201d our ant model of that. However general we\u2019ve been attempting to development in direction of fewer instruments. The final set of instruments we launched was the duty software, I feel \u2014 and we attempt to give Claude extra common variations to do issues.<\/p>\n<\/blockquote>\n<h4 id=\"what-s-the-latest-evolution-of-your-file-editing-tool-\">What\u2019s the newest evolution of your file modifying software?<\/h4>\n<p>29:03<\/p>\n<p>I&#8217;ve a long-running fascination with file modifying instruments\u2014they had been the topic of the previous Aider code modifying leaderboard, and I\u2019ve watched with curiosity as they\u2019ve developed in several coding brokers from search-and-replace based mostly to line-number-based to extra sophisticated patterns.<\/p>\n<p>The Claude API docs describe a textual content modifying software that\u2019s advisable for constructing in opposition to the API, however Claude Code appears to make use of barely completely different approaches right here.<\/p>\n<blockquote>\n<p>Simon: One of the crucial fascinating instruments is the file modifying software \u2014 you possibly can have file modifying as a software, or you possibly can inform it to make use of sed and grep and do issues that approach. What\u2019s the newest evolution of your file modifying software?<\/p>\n<p>Thariq: We nonetheless have one, however for instance we eliminated our grep and different search instruments \u2014 glob instruments \u2014 in favor of native bash. Like I stated in my discuss earlier, the fashions are sort of extra of a biology than a physics, and power design particularly is kind of arduous. I\u2019m undecided if Cat disagrees and thinks there\u2019s a science to the eval of it, however I feel software design is extra of an artwork, possibly \u2014 or a biology.<\/p>\n<p>Cat: I largely agree, however usually as we introduce extra instruments, we attempt to maintain the cardinality fairly low and ensure that each software we add has a definite operate from each different software, in order that Claude can very simply distinguish when to name every. For file edit, the rationale we&#8217;ve it is usually because we are able to render it. We present folks when Claude makes a file change, and there\u2019s this good devoted UI that claims: do you approve this edit to this file? The explanation we had a devoted file edit software was in order that we may deterministically know that Claude was making a file change, so we may present folks this good UI. A number of new customers onboarding nonetheless actually like this expertise, so we\u2019ve stored it round. However for lots of us who&#8217;re on auto mode proper now \u2014 hopefully you\u2019re not on YOLO mode \u2014 I don\u2019t assume it really issues, and we may in all probability simply take away file edit and be completely wonderful.<\/p>\n<\/blockquote>\n<h4 id=\"what-s-the-advice-within-anthropic-for-safely-running-claude-code-\">What\u2019s the recommendation inside Anthropic for safely operating Claude Code?<\/h4>\n<p>30:58<\/p>\n<p>It\u2019s the immediate injection query! Who higher than Anthropic workers to elucidate how Anthropic sees the danger of immediate injection assaults inflicting their Claude Code cases to run amok?<\/p>\n<p>It seems they actually belief their auto mode\u2014and see that because the characteristic that enabled Claude Tag.<\/p>\n<blockquote>\n<p>Simon: Let\u2019s speak about security and safety. I&#8217;m deeply conscious of the dangers of immediate injection, and there are such a lot of dangerous issues that may occur if any individual else tells my Claude Code what to do. I nonetheless principally run Claude Code in YOLO mode and really feel extremely responsible about it. What\u2019s the recommendation inside Anthropic for safely operating Claude Code?<\/p>\n<p>Cat: Why not auto mode?<\/p>\n<p>Simon: I&#8217;m beginning to use auto mode, however I don\u2019t perceive it sufficient to get how protected it&#8217;s. As of possibly three weeks in the past, I\u2019m defaulting to auto mode.<\/p>\n<p>Cat: Broadly inside Anthropic, nearly each single particular person makes use of auto mode. It&#8217;s the easiest way to do long-running work in Claude Code whereas being protected. We\u2019ve performed intensive bashing. Now we have 1000&#8217;s of evals. We\u2019ve commissioned many purple teamers to create adversarial environments with a view to trick Claude Code into doing dangerous actions, and we\u2019ve mitigated each single concern that they discovered. We\u2019re going to publish some evals within the coming weeks, however we\u2019ve just about mitigated each assault.<\/p>\n<p>Simon: That may be a huge declare.<\/p>\n<p>Cat: We\u2019ll share the evals for it so people can assess, however we\u2019ve been extraordinarily diligent about figuring out all of the methods during which Claude would possibly mess up after which updating auto mode to counter it. It doesn\u2019t catch 100% of issues \u2014 that might be approach too robust a declare. However for the primary classes of dangers that we\u2019re involved about, like immediate injection and information exfiltration, the dangers are far decrease than the typical human reviewer.<\/p>\n<\/blockquote>\n<p>I&#8217;m very a lot trying ahead to studying extra about their evals and strategy to verifying auto mode.<\/p>\n<blockquote>\n<p>Thariq: A bit of on how auto mode works \u2014 it\u2019s helpful to construct this psychological mannequin. At any time when Claude is doing a flip, or a bash name, there\u2019s a Sonnet classifier that&#8217;s judging the software name and in addition the context of the dialog \u2014 your instruction. There are some issues round permissions which can be dependent in your request: you don\u2019t wish to give git push permissions on a regular basis, however in the event you say \u201cpush this to GitHub,\u201d you need it to do it \u2014 and in the event you say \u201cdon\u2019t push,\u201d you need it to disclaim it. Auto mode will try this. That exact factor occurs to me quite a bit, the place Claude tried to do one thing as a result of it\u2019s very useful and proactive, and auto mode noticed \u201cdon\u2019t do that\u201d and surfaced it. So it\u2019s good on the dynamic permissions that you just your self give contained in the immediate, which I feel is admittedly vital. It additionally works properly with our sandboxing infrastructure, as a result of sandboxing is a kind of issues the place there are such a lot of completely different edge circumstances that it\u2019s arduous for us to deterministically comply with them. Now we have a sandbox, and when one thing wants to flee the sandbox \u2014 like a community request \u2014 auto mode can take a look at that request and ask: does this make sense? \u2014 and permit it.<\/p>\n<p>Simon: I hadn\u2019t realized auto mode is interacting with the networking sandbox as properly.<\/p>\n<p>Cat: It interacts with any permission immediate the consumer would in any other case see.<\/p>\n<p>Simon: How previous is auto mode? As a characteristic I had entry to, it\u2019s solely a few months previous, proper?<\/p>\n<\/blockquote>\n<p>(It was first made out there to the general public on March twenty fourth.)<\/p>\n<blockquote>\n<p>Cat: We\u2019ve been utilizing it inside Anthropic since January, so we\u2019ve been hardening it for fairly some time. Anthropic is extraordinarily targeted on security and safety, and we\u2019ve been working broadly throughout our alignment and safeguards groups to allow the rollout internally, construct out these evals, and make auto mode much more strong earlier than sharing it with the world.<\/p>\n<p>Thariq: That is additionally the rationale Claude Tag is so good \u2014 Claude Tag makes use of auto mode. I\u2019ve heard numerous build-versus-buy questions on a Slackbot, and I\u2019m like: please, you in all probability shouldn\u2019t construct your individual AI Slackbot. There are such a lot of assault vectors. You could have a suggestions channel that customers can publish suggestions into, and now your bot is studying it. The work we\u2019ve put in with auto mode \u2014 and we&#8217;ve a common Swiss cheese protection for safety; we additionally RL in opposition to these items \u2014 I feel that is actually what makes Claude Tag work. It really works seamlessly together with your permissions, and also you don\u2019t wish to be immediate injected in your Slack.<\/p>\n<\/blockquote>\n<h4 id=\"are-there-more-security-things-in-the-pipeline-beyond-auto-mode-\">Are there extra safety issues within the pipeline past auto mode?<\/h4>\n<p>35:54<\/p>\n<blockquote>\n<p>Simon: Are there any extra safety issues within the pipeline that transcend auto mode?<\/p>\n<p>Thariq: I feel we\u2019re very safe. With Claude Tag you possibly can provision your individual credentials for Claude, so it doesn\u2019t must act in your behalf \u2014 you possibly can have Claude as an identification, and that additionally makes it simpler to audit and examine what Claude is doing.<\/p>\n<p>Simon: As a result of Claude Tag is influenced by anybody who can discuss to it \u2014 it\u2019s bought a a lot wider pool of individuals telling it what to do.<\/p>\n<p>Thariq: That\u2019s proper. And naturally we&#8217;ve probes as properly with Fable, which is a downstream impact of our security and analysis work. I feel that is the second the place you see Anthropic being an AI security firm actually paying off: we actually need Claude to have the ability to run in an aligned approach over lengthy durations of time, and auto mode must be principally flawless for this to work \u2014 it\u2019s all downstream of our being an AI security firm.<\/p>\n<p>Cat: We additionally launched trusted gadgets for the distant management customers on the market who wish to be safer. And for all of our distant environments, we assist credential injection. If you need Claude Code to have the ability to entry Datadog, however you don\u2019t need Claude Code itself to carry the Datadog credential, you possibly can arrange our identification and credential administration system in order that the Datadog credentials are solely usable by the agent however not accessible by the agent \u2014 we insert them on the fly when the agent tries to make a Datadog request.<\/p>\n<\/blockquote>\n<p>I actually like that credential injection sample, the place Claude Code can entry an API through a proxy and that proxy each audits the request and injects the related API key\u2014so Claude can entry authenticated endpoints with out gaining access to the API credentials itself.<\/p>\n<h4 id=\"how-has-the-past-year-and-a-half-changed-how-you-think-about-your-own-craft-\">How has the previous 12 months and a half modified how you concentrate on your individual craft?<\/h4>\n<p>37:53<\/p>\n<p>Thariq talked a couple of sense of grief introduced on by Fable-class fashions in his keynote within the morning, and we dived additional into that as a part of our dialog. I\u2019ve been calling this Deep Blue.<\/p>\n<blockquote>\n<p>Simon: Let\u2019s discuss a bit of bit concerning the human component. Lots of people are feeling a way of loss now that a lot of what they thought of to be their function in constructing software program is being subsumed by the fashions. How do you concentrate on that? How has the previous 12 months and a half modified the way in which you concentrate on your individual craft and the worth that you just add?<\/p>\n<p>Thariq: Cat and Boris are such good reminders that it&#8217;s important to be extra bold. They\u2019re at all times like: we\u2019re rising so quick, we&#8217;ve to be on the sting, we&#8217;ve to do the most effective work we are able to. That\u2019s a relentless reminder for me \u2014 any time I\u2019m sluggish on one thing, I\u2019m like, okay, can I do it sooner? Can I be extra bold right here? And oftentimes the reply is Claude, as a result of Claude is getting higher as you go \u2014 the final time I attempted this, it was with the earlier mannequin. In your level about loss: I feel that is actual. In the event you\u2019re solely attempting to do the identical work you had been doing earlier than LLMs, and now it\u2019s a immediate, it&#8217;s, I feel, sort of a tragic feeling. And the way in which you offset that&#8217;s by being extra bold. I feel Jared is such  instance \u2014 he hand-wrote all the Zig code in his Oakland house in a couple of 12 months, barely left his home, and had a lot enjoyable doing that. Now I see him rewrite all of Bun into Rust and he\u2019s having a lot enjoyable doing that \u2014 it\u2019s a lot extra bold, and that\u2019s how he offsets it. Typically it\u2019s asking how do I do the larger factor and do extra \u2014 I feel success is enjoyable. It\u2019s altering your ambition.<\/p>\n<\/blockquote>\n<p>\u201cThe best way you offset that&#8217;s by being extra bold\u201d neatly captures the place I\u2019ve landed on this concern myself as properly.<\/p>\n<blockquote>\n<p>Simon: And Cat, what does that appear like from a product administration perspective?<\/p>\n<p>Cat: I really feel just like the product function simply modifications each single month. All of the PMs on our crew are this mixture of engineer, designer, PM \u2014 most of them really was once full-time engineers. For us it actually means plugging in at any time when there\u2019s any sort of hole. If we&#8217;ve an thought and we didn\u2019t encourage any engineer to go construct it, then we must always simply construct it, put it right into a pocket book, and encourage folks to take it to manufacturing. If the designs look a bit of off, let\u2019s take a web page that\u2019s comparable, do a first-pass design, and tag in somebody who\u2019s very detail-oriented to fill within the gaps. Or if we discover that our crew and product adoption is larger inside the firm, and extra folks must know what\u2019s coming down the pipe for Claude Code, Claude Tag, and Cowork \u2014 let\u2019s automate determining our complete launch calendar, let\u2019s automate getting these standing updates asynchronously so we\u2019re not bugging folks, and ensure our updates in our inner announce channels are totally detailed and to the purpose. For us it\u2019s very a lot understanding what the hole is true now between a terrific thought and getting one thing to our clients, and the way will we automate it as a lot as potential.<\/p>\n<\/blockquote>\n<p>This displays one thing I\u2019ve seen: when you possibly can produce code a lot sooner, time spent blocked awaiting a choice from another person turns into a way more notable bottleneck. Engineers who could make product selections can transfer a complete lot sooner, and the price of getting a kind of selections unsuitable is far much less prohibitive.<\/p>\n<h4 id=\"what-s-a-moment-when-claude-has-surprised-you-\">What\u2019s a second when Claude has shocked you?<\/h4>\n<p>41:50<\/p>\n<blockquote>\n<p>Simon: What\u2019s a second when Claude has shocked you? When the mannequin did one thing you didn\u2019t assume it could be capable to do?<\/p>\n<p>Thariq: I\u2019ve posted quite a bit about Claude video modifying, however most just lately I gave a chat on the ACM Agentic convention, and I requested, &#8220;Hey guys, do you might have the edited video? I\u2019d like to publish it and share it with my comms crew.&#8221; They stated, &#8220;Oh, it\u2019s taking so lengthy.&#8221; So I requested for the uncooked recordsdata. They despatched me the video of me speaking on stage, the video of the deck, and the audio file, and stated, &#8220;Good luck.&#8221; I gave this to Claude, together with my HTML deck, and stated, &#8220;Hey, are you able to simply edit this collectively?\u201c And what it does is truthfully unimaginable \u2014 I\u2019m able to ship it. It transcribes all the video. It notices that typically the video of my deck is a bit of bizarre \u2014 there\u2019s a popup of an auto-update within the center \u2014 and it goes, \u201dOh, I in all probability shouldn\u2019t use the video of your deck. What I\u2019m going to do is slice it up, determine which slide you\u2019re on, and use the HTML supply as an alternative.&#8221; So it shows the HTML supply. Then it\u2019s bought video of me, however I\u2019m solely taking on a small a part of the stage, so it\u2019s cropping dynamically to the place I&#8217;m on the stage \u2014 and I\u2019m pacing, so it\u2019s monitoring me as I tempo. And it\u2019s transcribing what I\u2019m saying.<\/p>\n<p>Simon: This was Fable, proper?<\/p>\n<p>Thariq: This was Fable, yeah. It was  immediate, however it was a one-shot immediate. Then I requested it so as to add some fascinating animations and graphics, and I used to be simply blown away. It does ffmpeg, it does Remotion.<\/p>\n<\/blockquote>\n<p>Right here\u2019s Thariq\u2019s video on how he used Fable to edit Fable\u2019s personal launch video, and right here\u2019s that launch video.<\/p>\n<h4 id=\"what-can-t-it-do-yet-\">What can\u2019t it do but?<\/h4>\n<p>43:36<\/p>\n<p>I\u2019m embarrased to confess that I\u2019ve been discovering it fairly arduous to provide you with duties that frontier fashions like Fable 5 and GPT-5.6 are unable to perform.<\/p>\n<p>Cat nonetheless doesn\u2019t price its UX design expertise:<\/p>\n<blockquote>\n<p>Simon: What can\u2019t it do? What are the issues the place you\u2019re nonetheless dissatisfied \u2014 the place you\u2019re ready for Claude Fable 6 to determine it out for you?<\/p>\n<p>Cat: I would like it to have higher design and UX style. It\u2019s now on the level the place if I write out a immediate with an in depth spec of how I need a characteristic to behave, it&#8217;ll often behave that approach. However the paddings is likely to be off, or the interface simply isn\u2019t pleasant but. It leans on present finest practices for the way apps are designed, however for frontier AI merchandise, there are such a lot of new interplay experiences that we&#8217;ve but to design.<\/p>\n<p>Simon: There\u2019s an Opus aesthetic \u2014 you possibly can take a look at one thing and go, \u201cYeah, that was designed by Opus.\u201d It\u2019d be good if we may transfer past that.<\/p>\n<p>Cat: Yeah. I\u2019m very excited for future fashions to hopefully be interplay design thought companions.<\/p>\n<p>Thariq: What can\u2019t it do? I might like to see it work together extra with the true world. Can it resolve science? Can it orchestrate the experiments? There\u2019s some quantity of coding that goes into that, however there\u2019s additionally this different style of the broader world that it wants.<\/p>\n<\/blockquote>\n<h4 id=\"which-parts-of-anthropic-s-culture-should-other-companies-steal-\">Which elements of Anthropic\u2019s tradition ought to different corporations steal?<\/h4>\n<p>45:11<\/p>\n<p>I figured this is able to make a terrific closing query:<\/p>\n<blockquote>\n<p>Simon: Which elements of Anthropic\u2019s firm tradition do you assume uniquely assist Anthropic be productive with these instruments, that different corporations ought to steal? What are the cultural hacks folks needs to be adopting from you?<\/p>\n<p>Cat: I\u2019ll share one for Claude Tag. Claude Tag works finest when you might have it in a public channel, and when most of your channels are public. Claude Tag is ready to search throughout all public channels to get as a lot context as potential to provide the highest-accuracy reply \u2014 and it\u2019s solely ready to do that if it has entry to every little thing.<\/p>\n<p>Thariq: I discussed this in my keynote, however it\u2019s so vital to me I wish to re-emphasize it. The co-founders say we don\u2019t negotiate in opposition to ourselves, and I feel that is actually vital. You&#8217;ll be able to think about trade-offs in your head and discuss your self out of doing one thing bold \u2014 or you possibly can simply attempt to do the bold factor. We\u2019re so usually asking: what if we simply did it? Is that this an actual trade-off or not? And if that&#8217;s the case, why \u2014 the place\u2019s the proof that it\u2019s an actual trade-off, and never simply one thing that sounds cheap? Make the trade-offs present themselves to you. Be as bold as you possibly can.<\/p>\n<\/blockquote>\n<h4 id=\"what-s-your-favorite-absurd-thing-you-ve-built-with-claude-just-because-you-could-\">What\u2019s your favourite absurd factor you\u2019ve constructed with Claude, simply since you may?<\/h4>\n<p>46:46<\/p>\n<p>I couldn\u2019t resist throwing on this one as properly.<\/p>\n<blockquote>\n<p>Simon: What\u2019s one in every of your favourite absurd issues that you just\u2019ve constructed with Claude, simply since you may construct it?<\/p>\n<p>Thariq: I\u2019m engaged on a 2D Avenue Fighter preventing sport with me as a personality \u2014 and my pals as properly. It makes use of Claude Code to immediate Gemini \u2014 and truthfully the Seedance mannequin is fairly good \u2014 to make video animations. It really works nice; it\u2019s so good at prompting, and it could actually confirm the frames to test whether or not an animation was good.<\/p>\n<p>Simon: Is that this Avenue Fighter 2-level 2D sprites you\u2019re producing?<\/p>\n<p>Thariq: Yeah, precisely \u2014 2D sprites. The animation appears wonderful. And it could actually additionally determine hitboxes \u2014 it may be like, \u201cOh, your fist is right here, I\u2019ll draw the JSON hitbox.\u201d It\u2019s unimaginable.<\/p>\n<p>Cat: Mine is rather more easy. I\u2019m an enormous rock climber and numerous my pals climb, so we&#8217;ve this little app we constructed with Claude Code the place we log all of the initiatives we\u2019re engaged on. We additionally go outside collectively quite a bit, so we&#8217;ve Claude do all this analysis with workflows. Workflows is wonderful \u2014 we model it as a coding software, however it\u2019s wonderful for doing deep analysis for journey. I additionally plan our crew offsites, and it\u2019s good at discovering venues that may match all of us. I exploit workflows to analysis all of the climbing locations we&#8217;d wish to go to, and what has direct flights from the place all of us are situated. It goes to Mountain Mission and finds all of the climbs at our grade degree. It finds the Airbnb. And I don\u2019t like mountain climbing, so I care quite a bit about it having a really brief strategy \u2014 very brief strolling distance from the place the automobile parks to the place the rock really is \u2014 and it filters for this. With present apps I&#8217;ve to manually click on by way of Mountain Mission, however with this I simply put in all of our preferences and it\u2019s a customized app for us.<\/p>\n<p>Simon: So that you\u2019re principally vibe coding Jira for mountaineering.<\/p>\n<p>Cat: Precisely.<\/p>\n<\/blockquote>\n<h4 id=\"audience-any-plans-for-eval-building-tools-and-agent-observability-\">Viewers: Any plans for eval-building instruments and agent observability?<\/h4>\n<p>49:23<\/p>\n<p>We had a couple of minutes on the finish for questions from the viewers.<\/p>\n<blockquote>\n<p>Viewers: Do you might have any near-term plans to construct extra eval instruments for us to construct eval datasets, and extra observability instruments to watch the efficiency of brokers and workflows?<\/p>\n<p>Cat: We\u2019ve thought of constructing eval instruments, however I feel the limiting issue really tends to be that it takes a very long time for patrons to construct actually high-quality evals. So I feel the tooling is much less of the constraint, and extra the ability set of the way you construct a terrific eval. That\u2019s an space the place we\u2019re excited to each make investments internally and hopefully share some finest practices externally.<\/p>\n<\/blockquote>\n<h4 id=\"audience-how-is-memory-designed-today-and-would-you-move-from-files-to-a-data-store-\">Viewers: How is reminiscence designed right this moment \u2014 and would you progress from recordsdata to a knowledge retailer?<\/h4>\n<p>50:08<\/p>\n<blockquote>\n<p>Viewers (Sai): I\u2019m  within the reminiscence and the multiplayer. How is reminiscence being designed right this moment? I assume it\u2019s round recordsdata. And second, have you considered an orthogonal route the place you&#8217;d really want an information retailer for these reminiscences, as an alternative of recordsdata, to scale it higher?<\/p>\n<p>Thariq: Proper now for Claude Tag the reminiscence is channel-specific. Each Claude in that channel has a shared reminiscence, and the cases have a session \u2014 however the session can contribute again to predominant reminiscence. We do numerous reminiscence analysis, and it may be sort of unintuitive what the fitting method to do reminiscence is. We\u2019re at all times operating reminiscence experiments. The way it works proper now in Claude Tag is a markdown file per channel.<\/p>\n<\/blockquote>\n<\/div>\n<p><br \/>\n<br \/><a href=\"https:\/\/simonwillison.net\/2026\/Jul\/21\/cat-and-thariq\/#atom-everything\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>A Fireplace Chat with Cat and Thariq from the Claude Code crew twenty first July 2026 Earlier this month I hosted a hearth chat session on the AI Engineer World\u2019s Truthful with Cat Wu and Thariq Shihipar from Anthropic\u2019s Claude Code crew. We talked about Claude Code, Claude Tag, Fable, coding agent safety, evals, software [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":2686,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"fifu_image_url":"https:\/\/img.youtube.com\/vi\/uU5Gv2h8-9g\/maxresdefault.jpg","fifu_image_alt":"","jnews-multi-image_gallery":[],"jnews_single_post":[],"jnews_primary_category":[],"jnews_override_bookmark_settings":[],"jnews_social_meta":[],"jnews_override_counter":[],"footnotes":""},"categories":[2],"tags":[3164,2632,182,362,3163,632,3165],"class_list":["post-2684","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ai-research-breakthroughs","tag-cat","tag-chat","tag-claude","tag-code","tag-fireside","tag-team","tag-thariq"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v27.7 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>A Fireplace Chat with Cat and Thariq from the Claude Code crew - Future News 24<\/title>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/futurenews24.com\/index.php\/2026\/07\/21\/atom-everything-31\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"A Fireplace Chat with Cat and Thariq from the Claude Code crew - Future News 24\" \/>\n<meta property=\"og:description\" content=\"A Fireplace Chat with Cat and Thariq from the Claude Code crew twenty first July 2026 Earlier this month I hosted a hearth chat session on the AI Engineer World\u2019s Truthful with Cat Wu and Thariq Shihipar from Anthropic\u2019s Claude Code crew. We talked about Claude Code, Claude Tag, Fable, coding agent safety, evals, software [&hellip;]\" \/>\n<meta property=\"og:url\" content=\"https:\/\/futurenews24.com\/index.php\/2026\/07\/21\/atom-everything-31\/\" \/>\n<meta property=\"og:site_name\" content=\"Future News 24\" \/>\n<meta property=\"article:published_time\" content=\"2026-07-21T12:54:00+00:00\" \/>\n<meta property=\"article:modified_time\" content=\"2026-07-22T06:59:40+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/img.youtube.com\/vi\/uU5Gv2h8-9g\/maxresdefault.jpg\" \/>\n<meta name=\"author\" content=\"Future News 24\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:image\" content=\"https:\/\/img.youtube.com\/vi\/uU5Gv2h8-9g\/maxresdefault.jpg\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"Future News 24\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"46 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/21\\\/atom-everything-31\\\/#article\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/21\\\/atom-everything-31\\\/\"},\"author\":{\"name\":\"Future News 24\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/person\\\/cecad1bde21cfc357cf70128144d6c83\"},\"headline\":\"A Fireplace Chat with Cat and Thariq from the Claude Code crew\",\"datePublished\":\"2026-07-21T12:54:00+00:00\",\"dateModified\":\"2026-07-22T06:59:40+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/21\\\/atom-everything-31\\\/\"},\"wordCount\":9182,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/21\\\/atom-everything-31\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/img.youtube.com\\\/vi\\\/uU5Gv2h8-9g\\\/maxresdefault.jpg\",\"keywords\":[\"Cat\",\"Chat\",\"Claude\",\"Code\",\"Fireside\",\"team\",\"Thariq\"],\"articleSection\":[\"AI Research &amp; Breakthroughs\"],\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/21\\\/atom-everything-31\\\/#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/21\\\/atom-everything-31\\\/\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/21\\\/atom-everything-31\\\/\",\"name\":\"A Fireplace Chat with Cat and Thariq from the Claude Code crew - Future News 24\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/21\\\/atom-everything-31\\\/#primaryimage\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/21\\\/atom-everything-31\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/img.youtube.com\\\/vi\\\/uU5Gv2h8-9g\\\/maxresdefault.jpg\",\"datePublished\":\"2026-07-21T12:54:00+00:00\",\"dateModified\":\"2026-07-22T06:59:40+00:00\",\"breadcrumb\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/21\\\/atom-everything-31\\\/#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/21\\\/atom-everything-31\\\/\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/21\\\/atom-everything-31\\\/#primaryimage\",\"url\":\"https:\\\/\\\/img.youtube.com\\\/vi\\\/uU5Gv2h8-9g\\\/maxresdefault.jpg\",\"contentUrl\":\"https:\\\/\\\/img.youtube.com\\\/vi\\\/uU5Gv2h8-9g\\\/maxresdefault.jpg\"},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/2026\\\/07\\\/21\\\/atom-everything-31\\\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/futurenews24.com\\\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"A Fireplace Chat with Cat and Thariq from the Claude Code crew\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#website\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/\",\"name\":\"Future News 24\",\"description\":\"The Smart Hub for AI and Next-Gen Innovation\",\"publisher\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\\\/\\\/futurenews24.com\\\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#organization\",\"name\":\"Future News 24\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/logo\\\/image\\\/\",\"url\":\"https:\\\/\\\/futurenews24.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/fn24-favicon.png\",\"contentUrl\":\"https:\\\/\\\/futurenews24.com\\\/wp-content\\\/uploads\\\/2026\\\/06\\\/fn24-favicon.png\",\"width\":250,\"height\":250,\"caption\":\"Future News 24\"},\"image\":{\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/logo\\\/image\\\/\"}},{\"@type\":\"Person\",\"@id\":\"https:\\\/\\\/futurenews24.com\\\/#\\\/schema\\\/person\\\/cecad1bde21cfc357cf70128144d6c83\",\"name\":\"Future News 24\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"url\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"contentUrl\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g\",\"caption\":\"Future News 24\"},\"sameAs\":[\"https:\\\/\\\/futurenews24.com\"],\"url\":\"https:\\\/\\\/futurenews24.com\\\/index.php\\\/author\\\/mridulpahuja20\\\/\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"A Fireplace Chat with Cat and Thariq from the Claude Code crew - Future News 24","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/futurenews24.com\/index.php\/2026\/07\/21\/atom-everything-31\/","og_locale":"en_US","og_type":"article","og_title":"A Fireplace Chat with Cat and Thariq from the Claude Code crew - Future News 24","og_description":"A Fireplace Chat with Cat and Thariq from the Claude Code crew twenty first July 2026 Earlier this month I hosted a hearth chat session on the AI Engineer World\u2019s Truthful with Cat Wu and Thariq Shihipar from Anthropic\u2019s Claude Code crew. We talked about Claude Code, Claude Tag, Fable, coding agent safety, evals, software [&hellip;]","og_url":"https:\/\/futurenews24.com\/index.php\/2026\/07\/21\/atom-everything-31\/","og_site_name":"Future News 24","article_published_time":"2026-07-21T12:54:00+00:00","article_modified_time":"2026-07-22T06:59:40+00:00","og_image":[{"url":"https:\/\/img.youtube.com\/vi\/uU5Gv2h8-9g\/maxresdefault.jpg","type":"","width":"","height":""}],"author":"Future News 24","twitter_card":"summary_large_image","twitter_image":"https:\/\/img.youtube.com\/vi\/uU5Gv2h8-9g\/maxresdefault.jpg","twitter_misc":{"Written by":"Future News 24","Est. reading time":"46 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/21\/atom-everything-31\/#article","isPartOf":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/21\/atom-everything-31\/"},"author":{"name":"Future News 24","@id":"https:\/\/futurenews24.com\/#\/schema\/person\/cecad1bde21cfc357cf70128144d6c83"},"headline":"A Fireplace Chat with Cat and Thariq from the Claude Code crew","datePublished":"2026-07-21T12:54:00+00:00","dateModified":"2026-07-22T06:59:40+00:00","mainEntityOfPage":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/21\/atom-everything-31\/"},"wordCount":9182,"commentCount":0,"publisher":{"@id":"https:\/\/futurenews24.com\/#organization"},"image":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/21\/atom-everything-31\/#primaryimage"},"thumbnailUrl":"https:\/\/img.youtube.com\/vi\/uU5Gv2h8-9g\/maxresdefault.jpg","keywords":["Cat","Chat","Claude","Code","Fireside","team","Thariq"],"articleSection":["AI Research &amp; Breakthroughs"],"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/futurenews24.com\/index.php\/2026\/07\/21\/atom-everything-31\/#respond"]}]},{"@type":"WebPage","@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/21\/atom-everything-31\/","url":"https:\/\/futurenews24.com\/index.php\/2026\/07\/21\/atom-everything-31\/","name":"A Fireplace Chat with Cat and Thariq from the Claude Code crew - Future News 24","isPartOf":{"@id":"https:\/\/futurenews24.com\/#website"},"primaryImageOfPage":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/21\/atom-everything-31\/#primaryimage"},"image":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/21\/atom-everything-31\/#primaryimage"},"thumbnailUrl":"https:\/\/img.youtube.com\/vi\/uU5Gv2h8-9g\/maxresdefault.jpg","datePublished":"2026-07-21T12:54:00+00:00","dateModified":"2026-07-22T06:59:40+00:00","breadcrumb":{"@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/21\/atom-everything-31\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/futurenews24.com\/index.php\/2026\/07\/21\/atom-everything-31\/"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/21\/atom-everything-31\/#primaryimage","url":"https:\/\/img.youtube.com\/vi\/uU5Gv2h8-9g\/maxresdefault.jpg","contentUrl":"https:\/\/img.youtube.com\/vi\/uU5Gv2h8-9g\/maxresdefault.jpg"},{"@type":"BreadcrumbList","@id":"https:\/\/futurenews24.com\/index.php\/2026\/07\/21\/atom-everything-31\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/futurenews24.com\/"},{"@type":"ListItem","position":2,"name":"A Fireplace Chat with Cat and Thariq from the Claude Code crew"}]},{"@type":"WebSite","@id":"https:\/\/futurenews24.com\/#website","url":"https:\/\/futurenews24.com\/","name":"Future News 24","description":"The Smart Hub for AI and Next-Gen Innovation","publisher":{"@id":"https:\/\/futurenews24.com\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/futurenews24.com\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/futurenews24.com\/#organization","name":"Future News 24","url":"https:\/\/futurenews24.com\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/futurenews24.com\/#\/schema\/logo\/image\/","url":"https:\/\/futurenews24.com\/wp-content\/uploads\/2026\/06\/fn24-favicon.png","contentUrl":"https:\/\/futurenews24.com\/wp-content\/uploads\/2026\/06\/fn24-favicon.png","width":250,"height":250,"caption":"Future News 24"},"image":{"@id":"https:\/\/futurenews24.com\/#\/schema\/logo\/image\/"}},{"@type":"Person","@id":"https:\/\/futurenews24.com\/#\/schema\/person\/cecad1bde21cfc357cf70128144d6c83","name":"Future News 24","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","url":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/d57f07142d73cb5503ab2446ea7bc9ef3d0a5ba378d64a6157692311e42bf097?s=96&d=mm&r=g","caption":"Future News 24"},"sameAs":["https:\/\/futurenews24.com"],"url":"https:\/\/futurenews24.com\/index.php\/author\/mridulpahuja20\/"}]}},"_links":{"self":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/2684","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/comments?post=2684"}],"version-history":[{"count":1,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/2684\/revisions"}],"predecessor-version":[{"id":2685,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/posts\/2684\/revisions\/2685"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/media\/2686"}],"wp:attachment":[{"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/media?parent=2684"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/categories?post=2684"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/futurenews24.com\/index.php\/wp-json\/wp\/v2\/tags?post=2684"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}