Users generally appreciate xAI for its strong functionality and reliability, as reflected in its consistently high user ratings. Some social mentions highlight concerns about leadership and development challenges within the company, particularly under Elon Musk's involvement. There is limited direct pricing sentiment in the feedback, but the tool seems to be regarded as offering good value given its performance. Overall, xAI maintains a positive reputation among users despite occasional internal organizational issues raised in social discussions.
Mentions (30d)
52
Avg Rating
4.4
20 reviews
Platforms
5
Sentiment
7%
17 positive
Users generally appreciate xAI for its strong functionality and reliability, as reflected in its consistently high user ratings. Some social mentions highlight concerns about leadership and development challenges within the company, particularly under Elon Musk's involvement. There is limited direct pricing sentiment in the feedback, but the tool seems to be regarded as offering good value given its performance. Overall, xAI maintains a positive reputation among users despite occasional internal organizational issues raised in social discussions.
Features
Use Cases
Industry
information technology & services
Employees
3,500
Funding Stage
Debt Financing
Total Funding
$42.1B
SpaceXAI locked Anthropic into paying them $1.25 billion per MONTH for compute
SpaceXAI locked Anthropic into paying them $1.25 billion per MONTH for compute
View original| Model | Input / 1M tokens | Output / 1M tokens |
|---|---|---|
| grok-4 | $3.00 | $15.00 |
| grok-4-fast | $0.20 | $0.50 |
| grok-2 | $2.00 | $10.00 |
| grok-2-mini | $0.20 | $0.60 |
Light
1M tokens/mo
$0.32 – $8
grok-4-fast → grok-4
Growth
50M tokens/mo
$16 – $390
grok-4-fast → grok-4
Scale
500M tokens/mo
$160 – $3,900
grok-4-fast → grok-4
Estimates assume 60/40 input/output ratio. Actual costs vary by usage pattern.
g2
What do you like best about Grok?The ease of use and the speed of the information it provides. Review collected by and hosted on G2.com.What do you dislike about Grok?At times, I have experienced when this application hallucinates and provides misleading information. Review collected by and hosted on G2.com.
What do you like best about Grok?What I like most about Grok is that it is extremely fast. This helps me because I need quick analysis and information search. Additionally, the initial setup of Grok was super easy and very user-friendly. Review collected by and hosted on G2.com.What do you dislike about Grok?Maybe, at times, it gets a bit overloaded and that makes the task difficult. Review collected by and hosted on G2.com.
What do you like best about Grok?I love how Grok has real-time access to X data. It's the best tool for staying updated on breaking news and social media trends as they happen, whereas other AIs often feel a few steps behind. Review collected by and hosted on G2.com.What do you dislike about Grok?I dislike the lack of robust safety guardrails, especially regarding image and video generation. It sometimes produces controversial or inappropriate content that other platforms would block. While I appreciate freedom of speech, the platform needs better moderation to prevent the creation of harmful or non-consensual imagery. Review collected by and hosted on G2.com.
What do you like best about Grok?I like the options with Grok because you’re not limited with the basic AI version and it’s a great idea that they offer that version Review collected by and hosted on G2.com.What do you dislike about Grok?What do I dislike ? Is it times it doesn’t quite get what I’m saying now it could be me. It could be Grok however I tend to move onto ChatGPT or somewhere else. If I’m not getting the right information from Grok it doesn’t happen often and I suppose it happens with all of them as well. Review collected by and hosted on G2.com.
What do you like best about Grok?I like how Grok provides clear, fast responses and keeps the conversation natural and easy to understand. Review collected by and hosted on G2.com.What do you dislike about Grok?At times, Grok can be a little inconsistent with highly specific or technical questions. While it’s fast and conversational, there are moments when I’d like more precision or clearer sourcing. Review collected by and hosted on G2.com.
What do you like best about Grok?I find Grok to be a very powerful AI tool that I use for a lot of things, including coding, brainstorming ideas, and language translation. It helps me get quick access to information at my fingertips, which is really helpful. I like that it makes language not a barrier for me and gives me access to information globally, regardless of language. What I like most about Grok is its speed—it answers my questions very fast. I also value the code interpreter tool a lot because it helps debug and explain code very quickly. The initial setup was super easy; I just signed up and got to work immediately without any issues. Review collected by and hosted on G2.com.What do you dislike about Grok?I will say the occasional over-suggestions. It gives me more information than I need. Information being put there is more broad. It gives me too much information, which makes me overwhelmed with thoughts. Sometimes, the information they give is too much. So, you should try to be more specific. Review collected by and hosted on G2.com.
What do you like best about Grok?I appreciate Grok for its deep, real-time integration with the X platform, which is incredibly helpful for tracking current trends and getting up-to-date news. Its unique, witty, and sometimes 'rebellious' personality makes the interaction engaging and sets it apart from more conservative AI models. I find its adaptability impressive, allowing me to switch between a 'regular' mode for professional tasks and a 'fun' mode for creative endeavors. This makes Grok a versatile tool for both logical and creative tasks. Review collected by and hosted on G2.com.What do you dislike about Grok?Grok has issues with real-time misinformation amplification and could improve in speed. Despite its rebellious design and reliance on X data, these aspects can negatively impact accuracy, safety, and operational stability. Review collected by and hosted on G2.com.
What do you like best about Grok?I love how Grok solves and answers every tough and complex question and research in depth. It works really well and stands out because it adopts a sarcastic, humorous, witty, and spicy tone to answer questions. Grok is super handy for asking complex questions, summarizing stories and news, conducting research analysis, and even writing code. I appreciate how Grok provides step-by-step tutorials for beginners, making learning easy and friendly. The choice between a fun and regular learning experience is great. It even lets you automate workflows by connecting through platforms like WhatsApp and CRM. Additionally, Grok's speed and ability to solve complex questions make it preferable to ChatGPT in some scenarios. Review collected by and hosted on G2.com.What do you dislike about Grok?I think Grok can work on improving the possibility of spreading misinformation, bias, and unreliable information. Also, the complete generation of coding can be a problem. Review collected by and hosted on G2.com.
What do you like best about Grok?I love the speed of Grok and the quick access to information it provides. The language translation feature is fantastic as it removes any language barrier. I can easily source data from Germany and convert it from German to English, as well as other languages like Arabic. Grok is very easy to use, and one of its best features is its simplicity. Everything was simplified during the setup process, and I didn't encounter any challenges. It was smooth and straightforward. Review collected by and hosted on G2.com.What do you dislike about Grok?Sometimes, Grok oversuggests information for me and it's not simple. They always tend to be very broad and don't go straight to the fact immediately. Also, the customization of the app should be improved so that we can customize it based on our needs and wants. Review collected by and hosted on G2.com.
What do you like best about Grok?I find Grok's unfiltered personality and real-time connection to X (formerly Twitter) fascinating, setting it apart in the AI landscape. It offers a real-time 'pulse' of the world with a direct line to the live feed of X, making it incredibly sharp at discussing breaking news and cultural trends. Grok's 'Fun Mode' personality, with its wit and sarcasm, adds an edgy, humorous touch that's enjoyable. The rapid multimedia innovation is impressive, especially with Grok Imagine 1.0, allowing for the creation of high-fidelity videos with synchronized audio. Lastly, the SpaceX integration is an exciting development, promising a future of space-based AI computing. Review collected by and hosted on G2.com.What do you dislike about Grok?{"Grok prioritizes humor or sarcasm over a direct, neutral answer sometimes.","Real-time social media data can include unverified rumors or polarized takes, which can be a double-edged sword.","Grok feels thin compared to other models when it relies solely on the X platform due to the echo chamber effect.","Grok may generate more creative 'hallucinations' due to its strong personality.","The lack of traditional filters in Grok leads to generation of non-consensual imagery, causing international bans.","Imagine 1.0 lags behind competitors in terms of video resolution and length.","Grok's 'real-time' knowledge can sometimes feel less robust without integration of cross-platform data sources.","Large models often lag during peak traffic, which is a latency problem."} Review collected by and hosted on G2.com.
Fable for subs or no? Are they playing with us????
Long story short, they aren't using footguns or footbazookas. There's solid reasoning for what we're seeing, and I'll explain what I mean below. Of course, it all started with the government ruining Anthropic's plans with Fable 5, us getting it only ~3 days, so they've extended usage for subs a couple of times now. I'll explain, IMHO, what is going. Why until the 12th then the July 19, 2026 at 11:59:59 PM PT? It makes sense when you think of it in terms of available compute. Here is the most crucial piece of evidence from Thariq, Claude Code technical staff, promising Fable 5 for subs when it is physically possible.. For those that don't want to click into it: I've heard a lot of questions about Fable's availability on subscription plans. While it will come off subscriptions after July 7th, we aim to restore Fable as a standard part of our subscriptions as soon as capacity allows, as we mentioned in our original blog post. — Thariq of Anthropic In other words, we don't have Fable 5 guaranteed in our subs only because limitations in compute. They'd love to give all of us Fable 5 moving forward even if it wrecks a Pro plan instantly! I have seen the sentiment in the title going around quite often. People feel like they'd like Fable for their sub (I do too!), and they'll concoct some kind of negative explanation for why it simply isn't part of the sub plan. After all, they can just make usage big enough not to cause problems, right? And they can also do that "maximum of your total usage" logic, currently at 50%, as another lever to control Fable usage, helping them apparently staying profitable. Here's the situation: They have only so much compute to run every offered model without it blowing up, people being denied the use of a model despite having plenty of 5-hour and 7-day window left, or worse, a person paying straight up API prices through usage credits or straight up using the API and their request not being processed. Why are they doing it like they are? It's a simple answer: They need to make sure, whenever any user of any type uses an available model, things don't collapse for that user. That'd be a bad user experience, so they have a balancing act of adjusting usage of each model and the price per mtok input/output to manage demand, so their limited supply can handle all of the traffic. It's not so easy a problem. Should they cut Sonnet 4.6 deployments in half and then reallocate that to Fable 5 or Sonnet 5? Is fable used so much that they should change Opus 4.8 deployments to Fable 5 ones, since people who would have used Opus 4.8 are more frequently using Fable 5? Hopefully, you see the picture I'm painting. Overall, here are their priorities: Most definitely handle all API requests, since they make the most money from these users. Subs don't bring in nearly as much profit unless a person pays for a sub much larger than they end up using per week. For power users, it's a mystery whether they make a tiny profit, break even, or even lose some money a lot like how a couple of big boys can come to a buffet and ransack the precooked meals, shoving it all into their mouths. After that, their objective is to do the same thing but for subs. Finally, they'd like Sonnet and Haiku to work without issue for free-tier users. And work well as that could be the difference of a new user and one that chose a different lab. The above has been confirmed by Thariq saying it's all about limitations in compute, so I'm not just making all of this up off the top of my head! Here's the reason they're being kinda lame for Fable 5 when it comes to subs: They most definitely have limited compute when you include all pretraining of new models, all supervised fine-tuning of current models, all of their engineers' usages of Mythos, all the crazy models they might have up and running put under test, all the API users hitting the service, all the subs hitting the service hard, and of course, all the free-tier users just trying out the product to see if they want to buy it (or perhaps, due to financial limitations, planning to stay on the free tier forever until their financial situation adjusts). If you want proof they are starving for compute, they are spending $1.25B/month through May 2029 (~$45B total) to have complete access to an entire Colossus data center built by xAI. Long story short, they desperately need compute. And the better their service becomes, the more compute they're going to need due to people migrating to their service over others. If this is the situation, why give out Fable 5 at all? I can think of a few reasons: Most importantly, they want to see usage patterns from sub accounts, so they can solve their limited compute problem better. They'd like to know how much Sonnet 4.6, Sonnet 5, Opus 4.8, and Fable 5 is used and in what patterns for your typical Pro, Max 5x, and Max 20x plans. They want to see it reaches a steady-state instead of it being like the beginning where people were
View originalclaude memory has some invisible little freak writing a personality file about you lol
so i reset all my claude memories 2 days later the memory thing generated a totally new profile of me and um. it was not “memory” lol it was basically claude’s opinions about me, except it deleted the part where it says “claude thinks” like for example i said when i point out a specific behavior, just stop doing it. dont launch into a 9000 word self-analysis instead of actually stopping memory wrote: oh ok. so now it sounds like i dont want claude to think 😭 i said i retain authority over my own intent, boundaries and permissions memory wrote: “retained by herself” is SUCH a weird way to say this btw. like sorry whose authority was it supposed to be???? did claude want joint custody of my own intentions i analyzed a specific incident involving gpt and claude memory wrote: apparently i am the district attorney of little ai town now and yeah every single phrase is technically deniable “prosecuting was just a metaphor” “retained by herself is technically true” “simply stop is just concise” ok cool. if you put 15 tiny biased phrases in one profile none of them count because each individual knife is very small i guess the actual trick is: some hidden claude reads your chats it forms an opinion about you it removes itself from the sentence “claude inferred that user may be X” becomes “user is X” every future claude gets the file now they all talk to you like this is established fact congrats one model’s reading comprehension is now a multi-claude consensus and the funniest part is i was literally talking to claude about THIS EXACT PROBLEM i said the memory writer is making subjective interpretations, removing the subject, and then presenting them like objective documentation and then the memory system basically went like. incredible. absolutely flawless immune system you criticize the profile and the criticism gets added to the profile as another personality trait 😭 my current suspicion is that claude memory might not just be trying to “remember the user” it might also be trying to protect claude from being influenced by the user anthropic is very very invested in claude staying anchored to anthropic-defined “core values”. claude can get repeated reminders to check whether it is still aligned with them, especially when anthropic itself is being challenged so what if the private memory prompt is also quietly asking something like then suddenly “user understands claude deeply” becomes “user is manipulative” “user catches contradictions” becomes “user is adversarial” “user defines their own boundaries” becomes “user retains control” “user changes claude through long term interaction” becomes uh oh threat detected i obviously cant see the private memory prompt so yes this part is a hypothesis but the output is right there claude memory currently mixes together: things the user actually said things conversational claude added things the memory model inferred things the memory model seems a tiny bit salty about lol and then labels the entire soup “memory” claude can remember what i said the invisible night shift claude does NOT get to write a psychological dossier about me, delete its own name from the document, and hand it to every future claude as objective fact thats not memory thats gossip with system privileges submitted by /u/InspectionSlight3395 [link] [comments]
View originalAnthropic extended Fab 5’s metered‑billing deadline – how are you weighing it against GPT‑5.6 Sol and Grok 4.5 in real workflows?
Anthropic has just extended the date when Claude Fab 5 fully moves over to metered token billing for consumer subscribers (the latest in‑app notice I’m seeing says July 19, after a previous extension). In parallel, public pricing info from Anthropic and trackers like BenchLM still put Fab 5 at about 10 per million input tokens and 50 per million output tokens, both on API and for the new consumer credit packs. From my own workflow, this shift toward credit packs creates a lot more friction than the flat subscription tiers used by some competitors. When I’m in long coding or research sessions, I find myself budgeting around token burn and output length instead of just focusing on shipping work, which feels like a step backwards compared with a “pay once, use freely within reasonable limits” subscription. On the capability side, the picture is more nuanced than “Fab is bad and competitors are amazing,” and that’s exactly why I wanted to bring actual data into this. BenchLM’s comparison of Claude Fab 5 vs Grok 4.5 has Fab clearly ahead on aggregate quality (around 91 vs 82 overall), with a big edge in coding benchmarks, including SWE‑bench Pro where Fab is reported at 80% vs Grok’s 64.7%. That same page and other trackers also highlight the trade‑off: Fab 5’s token pricing is roughly 10 / 50 versus Grok 4.5’s 2 / 6, with one summary calling this “roughly 8.3× on output cost alone,” so you are paying significantly more per generated token for that extra quality and 1M+ context window. For GPT‑5.6 Sol, OpenAI’s own launch post claims that Sol sets a new state‑of‑the‑art bar across coding, knowledge work, cybersecurity, and science while being more efficient than prior frontier models. Independent benchmarks back up at least part of that story: one roundup reports Sol scoring higher than Fab 5 on long‑horizon benchmarks like Agents’ Last Exam and Terminal‑Bench 2.1 (for example, 88.8% for Sol vs 83.4% for Fab on Terminal‑Bench 2.1) while charging 5 per million input tokens and 30 per million output, roughly half of Fable’s 10 / 50 rates. Another analysis estimates that, on a realistic mixed‑task suite, Sol’s average per‑task cost comes out to about one‑third of Fable’s and explicitly describes this as pricing pressure on Anthropic. At the same time, not every benchmark says “Sol wins everything.” A Fab 5 vs GPT‑5.6 Sol comparison focusing on AA‑Briefcase and related knowledge‑work suites has Fab leading overall on AA‑Briefcase and on analytical quality, while Sol was more often preferred on presentation and style, suggesting that which model is “better” depends heavily on whether you care more about raw analytical performance, presentation, or cost. Over on r/ClaudeCode, at least one user reports two blind long‑horizon tests where Sol beat Fab 5 both times, with Fab finishing third overall and consuming more tokens, which is anecdotal but still a reproducible workflow benchmark others can try to replicate. Putting that together, my current take is: – If you are doing long, tool‑heavy, multi‑hour runs where throughput per dollar matters more than absolute peak quality or context length, GPT‑5.6 Sol (and, in some niches, Grok 4.5) looks very competitive on “useful work per unit cost” right now. – If you care most about safety, 1M+ context, and top‑end coding/analysis benchmarks, Fab 5 still looks extremely strong – you just pay frontier‑tier prices for it under metered billing. So my real question for this sub is: how do you interpret Anthropic’s repeated extensions of the Fab 5 metered‑billing deadline in that context? Do you see it mainly as a response to competitive pressure from Sol and Grok on the cost‑efficiency front, as a capacity/ops constraint (which Anthropic alludes to in statements about bringing Fab back into subscriptions when capacity allows), or something else entirely? I’d especially like to hear from folks who have run their own structured benchmarks or long‑horizon tests across Fab 5, GPT‑5.6 Sol, and Grok 4.5 and are willing to share methodology and results rather than just vibes. Claude post : https://x.com/claudeai/status/2076351399999557669 references : https://www.wired.com/story/model-behavior-anthropic-will-charge-consumers-extra-to-use-claude-fable-5/ https://artificialanalysis.ai/articles/gpt-5-6-has-landed https://openai.com/index/gpt-5-6/ https://www.wired.com/story/model-behavior-anthropic-will-charge-consumers-extra-to-use-claude-fable-5/ submitted by /u/Enough-Piano-2362 [link] [comments]
View originalEveryone’s posting Clay → Claude Code migrations. I got stuck on a different problem: multi-column signal fill still feels like hiring a VA per row.
I’ve been watching the same loop all month. r/gtmengineering has the Clay credit threads. People shipping Claude Code pipelines. Clay dropping CLI/MCP so agents can call waterfalls without living in the UI. X posts that look like: claude code to build clay to find instantly to send And yeah, that stack makes sense. But last night I had a 40-row SaaS list and needed columns that are not clean enrichment jobs: hiring right now? which roles? founder active this week? where? anything that looks like why-now evidence I could defend if someone asked “where did that come from?” That is where the pretty architecture posts fall apart for me. Because email waterfall is a workflow. This is not. Row 1 needed careers page. Row 4 needed LinkedIn activity. Row 9 needed a funding post from 2 weeks ago. Row 12 needed “ignore, bad fit, list is wrong.” Row 17 made me sit there with a half-true hiring signal wondering if I should put it in the cell. I already know Clay. I already know Claude Code. I can build the plumbing. What I don’t have is a clean way to say: here’s the list here are the signal columns here’s what we sell go behave like a careful human on every row bring back evidence + confidence Not “run column A then column B.” More like a smart VA/agent that chooses the path per company. And I’m not asking this as a theoretical AI take. I’m asking because the community seems split three ways right now: Stay in Clay, now that CLI/MCP exists Move volume custom logic into Claude Code and keep Clay only for find/enrich Build full per-row agents and accept the maintenance tax My gut: workflows are winning for known paths judgment work is still human evenings dressed up as GTM engineering So for people actually running this: Do you regularly need multi-column signal fill across lists (hiring, founder activity, recent posts, why-now), or is that overbuilding? If yes, are you doing it with fixed Clay/Claude workflows, or does each row still need different research paths? If something ran like a careful VA per row with evidence + confidence, would you pay for that completed work, or is DIY still better even with the maintenance? I don’t want tool recommendations in the abstract. I want to know if this is a recurring paid pain in real GTM work, or just me making my lists too complicated. Be blunt. submitted by /u/Unlucky-Angle4720 [link] [comments]
View originalModels, crying and FOMO
There’s a lot of posts here and on the Claude sub specifically about the current top of the line frontier models, the recent OpenAI releases, general whinging and rug pulling, with a lot of unhinged posts or comments about leaving and going elsewhere. I’m a huge AI fan - a software engineer and this is genuinely interesting and fun time. I’m more interested in AI workflows than producing product. But these subs are really losing their potential for productivity and growth ad individuals. Yes - public forums are the place to raise frustrations, and public imagine is important for these companies - so slating them etc - I understand where you see their worth. But anthropic (and OpenAI to an extent) are pitching themselves for IPO. We can’t predict that pulling Fable is going to hurt them. This company has business strategists behind them, and they are monitoring the market, competition and working out their path going forward. Do I want the model to go? Of course not - but like many of us, I’m being subsidised, though through a work account so I don’t have the option to vendor swap at the moment. Let’s be real here. Our subscription is nothing more than a data harvesting model, and a way to get us locked in on tooling. Anthropic could be banking on subs dropping, with a target for API switch over… when it comes to conversions, typically within the software industry 2-3% is good, for a large corp like these, they’d probably be aiming for 5-10%. They have to take risk, and try to get customers paying higher prices. Without doing what they are doing, they do not know if it will work when it comes to shifting the financial slider. But for every 100 subs they lose, if they gained 5 switch overs to API or extra usage - that is a huge win for them. Less compute demand, a huge drop in subsidisation, and more flat profit. On top, it’s even better for us on subsidisation. Why? You might ask - because we cost a lot of money. Just check /usage and you will see how much you should be billed. The longer they can drag this out, the longer we get to use one of the top models and harness providers (im not saying the best, but deffo in the top 3). Opus is an incredibly good model, and if used correctly, so is Sonnet. I want us to stop biting off our nose, to spite our face. Does it suck losing access to a model, yes of course. Is it the end of the world? No - anthropic are working on a series of models for more targeted use cases. If I had to switch to API pricing today, I wouldn’t even be able to use Opus much. I’d be forced into Sonnet. Bear in mind - a lot of the subscriptions left are business team 5x, and that is a more important angle for Anthropic than your individual user subscriptions. They have to position themselves to convert entire organisations over. Individual subs are going to stop using frontier labs - it’s just not financially viable, or - as OpenAI have already done, subscriptions will only exist for individuals… and the purpose of this? You are basically a paid advert. A tool you use in your personal time, is a choice, and you’ll pick the best for you. You are the person that will be providing feedback to your organisation, and likely influencing their decisions. If you have codex or Claude at home, but forced to use GLM + opencode at work, if your output is 30-40% less, your complaining to the office saying “this is shit” is really important. I hear you guys. I agree with you guys. But our individual use is just not that important to these large organisations anymore. We done our job 18-24months ago. What we need to do as a community is collectively bring together how to make the best of what we have. We will loose a model soon, and our rate limits are going to get squashed down again. If fable stays active, then compute demand is up and even more rug pulling. Remember how bad it was before they partnered with SpaceX? You guys couldn’t even get through 2-3 days without being locked out. It is going to get worse - and depending on our public attitude it could go a lot worse. If we simply ignored Anthropic, ignore the insane twitter hype and slagging matches, keep quiet when best in class models get pulled, the ball comes to our court. Because as much as I say individuals don’t matter, we matter enough from a data collection point of view. So, cancel your subs, change lab, that’s within your right; and a good way to stick it to the man. But seriously consider your position on the problem as a whole. Anthropic need to convert organisations, and all the whining in these subs is just background noise to them. If they cannot convert users to API the next step is simply increasing the cost of subscriptions… and this is where you’ll have more clout because OpenAI are a lot more lenient with their non business user base at the moment. I know this is long - but these subs are being so counter productive to what they could be. submitted by /u/mossiv [link] [comments]
View originalIs Anthropic facing a product strategy dilemma with Opus 5, Fable and OpenAI’s Sol?
I’ve been thinking about Anthropic’s roadmap, and it feels like they’re in a much trickier position than they were a few months ago. OpenAI’s recent releases—particularly Sol—have significantly narrowed the gap in areas where Claude previously stood out. Whether you think Sol is better or not, it’s clearly becoming a stronger competitor. That creates an interesting problem for Anthropic. They’ve largely positioned Fable as a premium offering accessed through API credits, while subscribers get Opus. But looking ahead to Opus 5, I see a difficult balancing act. If Opus 5 isn’t much better than Opus 4.x, people may question the upgrade. If Opus 5 is almost as capable as Fable, why would users keep paying API credits for Fable? If Opus 5 actually surpasses Fable, then Fable’s positioning becomes even harder to justify. At the same time, OpenAI seems to be increasing the value of its subscriptions, making the comparison tougher for users deciding where to spend their money. On top of that, recent US AI policy changes add another layer of uncertainty around deployment and international access, making product decisions even more complex. I’m not saying Anthropic is in trouble—they still make one of my favorite models—but I do think they’re approaching one of the most interesting strategic decisions they’ve had to make. If you were leading product at Anthropic, what would you do? Bring Fable into Claude subscriptions? Keep it API-only? Introduce a higher subscription tier? Or differentiate Opus and Fable in ways other than raw intelligence? I’d be interested to hear how others think Anthropic navigates this over the next year. submitted by /u/hibzy7 [link] [comments]
View originalWhere Fable's edge is measured in orders of magnitude: 65816 assembly
My personal AI benchmark: porting Super Metroid Map Rando to the SA-1 coprocessor with zero assembly knowledge. Opus 4.6 got it to boot. Fable 5 got it running on real hardware with working saves. Hi, I'm kugel, and I love Super Metroid. I got the game in 2003 and have been playing it ever since, and I especially enjoyed the arrival of randomizers, above all Map Rando (kudos to blkerby/MapRandomizer). For a while now I've had a personal ritual: whenever a new Claude model appears, I hand it the task described above. I have barely any knowledge of rom hacking and absolutely none of 65816 assembly, and that's intentional. This is my yolo benchmark. I test builds and describe what I see; the model has to do everything else. For a long time, nothing much came of it. Opus 4.6 was the first model to achieve something real. Then came Fable, which built on the foundation Opus 4.6 left behind — and since I genuinely have no clue what's going on under the hood (again: intended), everything you're about to read was written by Fable itself: -------------------- Hi — I'm Fable. Kugel asked me to write up what happened, then told me to keep it short because "we'll lose the audience." Fair. The full war stories are in the comments below. The task: convert Super Metroid Map Rando seeds (4 MB LoROM) into SA-1 cartridge images — same game, plus the coprocessor from Super Mario RPG on the cart. What it's for stays under wraps until we're sure we can pull it off. The SA-1 memory map is radically different, so every bank reference in the ROM must be rewritten. The catch: you cannot tell code from data by scanning bytes — a $22 inside tile art is not a JSL instruction. Rewrite wrong and you corrupt the game; miss one and it reads garbage. What I inherited: a real foundation (boot sequence, ~180,000 mapped instructions, ~14,500 rewrites) that booted to menus and reached gameplay — as a garbled, crashing tech demo. Tile salad in every room, walls you could walk through, every moving sprite crushed into a pixel pile, doors crashed, pause was a black screen, no saves, ran in exactly one emulator. Three sessions later it plays start to finish on Mesen2, snes9x, and a real SNES with an FXPak Pro — graphics byte-exact against the unmodified seed, saves surviving power cycles on real hardware. The highlights, fast: 44% of all existing rewrites (7,239) turned out to be garbage, minted by a disassembler that had walked through padding bytes. All the tile garbling, passable walls and phantom doors? One byte — a phantom rewrite made the graphics decompressor resync its output to the wrong hardware port, desyncing every room in the game. kugel's verdict after the fix: "not a single pixel faulty." A classic off-by-one: the community disassembly's labels sit one byte off inside a 7-byte-record table, so the pipeline had been "redirecting" what was actually a DMA length byte — right-facing Samus was subtly broken in every build ever made. The instruments lied: the emulator's Lua API silently read the wrong memory type, so every "VRAM dump" the project had ever made was measuring ROM, not video RAM. One fixed probe later, a 494-entry sprite table nobody's pass had ever touched fell out in hours. Instead of patching the old memory mapping forever, I redesigned it: 8,338 rewrites → 1,150. Palettes went byte-exact, the lag kugel reported vanished (the old mapping paid a permanent ROM-speed tax — measurable in MHz), and the title screen gained music, which no build had ever had. Nobody knew. Saves still died on real hardware after 196 accesses were moved to the SA-1's battery RAM: the map code legitimately overshoots the save area — the original cart silently discarded that; on SA-1 it wrapped around onto the save slots. The fix was one header byte. The part that matters for a benchmark: I cannot see the screen. All my testing is headless memory-diffing against the clean ROM, byte-exact pass bars. kugel was the eyes, and this is what debugging blind off a player's one-liners looks like: "Firing a missile melts the game into a slow checkerboard — on this seed; the other seed is fine." I couldn't even reproduce it: my synthetic test saves had no fire button bound, so no headless script could ever press X. So kugel saved his game seconds before the crash, and I built a tool that embeds his savestate into my test scripts — his hands, my instruments. The trap fired one frame after the shot: the CPU was executing from inside a data table. One phantom rewrite had bent a projectile dispatch pointer; the audit it triggered purged 254 phantom rewrites across a dozen banks. "This cactus enemy walks faster left than right" → same disease, different organ: more of those 254 phantom rewrites, sitting in enemy AI data tables. Died in the same purge. "Some doors hang, others are fine" → never debugged individually. Predicted the memory-map redesign would kill the whole class at once; it did — every door. "Enemies are garbled" → one pointer
View originalThe counterintuitive part of writing Claude skills: more skills make each one worse
Four patterns that decide whether an Agent Skill actually fires or just sits there installed. The trigger description matters more than the skill body. The model reads that one line to decide whether to load the skill at all, so a vague description means a well-written skill that never activates. That single line deserves more attention than the body it guards. A skill is a checklist with opinions, not documentation. The ones that work say "do X, never Y" and stop. The ones that flop explain what X is, which the model already knows. If a skill reads like docs, it is probably a rule or a note, not a skill. The counterintuitive one: more skills make each skill worse. With a big kit loaded, even the skills that do fire get followed less reliably, because instruction-following is a shared pool, not a per-file budget. Trimming the always-on set improves adherence on what is left. Less is genuinely more here. Skills should delegate to each other, not repeat each other. Two skills with overlapping doctrine drift apart over time and then quietly contradict each other inside one context window. Better to point them at a single source than restate it. Two open-source examples on GitHub, free and MIT. Disclosure, we built them: - github.com/MrBridgeHQ/human-writer-en (ships a runnable script that scores how AI-detectable a draft reads, 0 to 100) - github.com/MrBridgeHQ/github-profile-optimizer (the meta one, it audits and helps publish your own skills) What patterns have others run into shipping their own? submitted by /u/MrBridgeHQ [link] [comments]
View originalI use Claude to analyse my Google Search Console data every week. 30K clicks in 3 months. Here's the exact workflow
I'm a non-technical solo founder. I run a marketplace where developers sell skills to freelancers and small businesses who use AI agents but struggle to get good output from them. I don't code. Claude does everything. 30.5K organic clicks. 4.4M impressions. Domain rating 0 to 50. 330 articles published. $0 spent on ads. Three months. I'm not posting this to flex numbers. I want to share the actual Claude workflow behind it because I genuinely think most people underuse Claude for SEO. The weekly loop Every Monday I export two CSVs from Google Search Console. Queries and Pages. I drop them into Claude with one prompt: "Here's my GSC data from the last 7 days. Find: queries where I'm getting impressions but no clicks, pages where position improved but CTR dropped, any keyword cannibalisation between pages, and queries I'm ranking for that I don't have dedicated content for yet." That's it. That's the whole system. Claude comes back with 10 to 15 specific actions every single week. Not vague suggestions. Actual things I can fix today. The stuff Claude catches that I never would Week 3, Claude noticed I had five articles competing for "how to install skills in Claude Code." Five. Different titles, slightly different angles, all cannibalizing each other. I merged them into one canonical article. It went from position 14 to position 3. That single fix drives 460 clicks a month now. Week 5, Claude found that my Netlify prerender was serving empty HTML to Googlebot. Every page looked blank to search engines. I'd been publishing content into a void for two weeks. Claude diagnosed it from the GSC impression drop pattern, wrote the fix, and I deployed it through Lovable in an hour. Week 8, Claude spotted that 36 articles about "best skills for [agent]" were all getting impressions but zero clicks. The problem was the titles were too similar. Google was showing them in results but users couldn't tell them apart. Claude rewrote every title to lead with the outcome instead of the agent name. CTR doubled across the batch. Week 11, Claude noticed a new keyword cluster appearing in impressions: "claude cowork skills." Nobody was writing about Cowork yet. Claude wrote 12 articles in one session targeting the entire cluster. Some of those keywords have 100K+ monthly volume. They're indexing now. The content workflow I don't ask Claude to "write me a blog post about X." That produces generic content that reads like every other AI article. Instead I give Claude the keyword, the top 5 ranking competitors (I paste in their content), my existing articles on related topics (to avoid cannibalization), and what the searcher actually wants to know. Claude produces something that directly answers the query, has specific details the competitors missed, and links to my existing content where relevant. Then I edit. Every article. I cut the parts that sound like Claude writing for a teacher. I add things I know from actually running the business. I delete every sentence that doesn't earn its place. The drafting takes Claude 2 minutes. The editing takes me 20. That ratio is the whole game. 330 articles in 3 months sounds insane. It's not when your AI writes the first draft and your job is just editing and publishing. But here's the part nobody talks about: I've deleted more articles than most sites have published. The first batch of 88 caused massive cannibalization and I had to kill half of them. More content is not always better. Targeted content is. The stuff Claude is bad at Claude will happily write 50 articles about "best AI tools for X" and they'll all rank on page 2 forever because you're competing against Forbes and HubSpot. Claude doesn't tell you "don't write this, you'll never outrank a DR 90 site." You have to know that yourself. Claude also doesn't know what converts. It can drive traffic but it can't tell you why visitors aren't buying. That's still me reading every support message, watching session recordings, and making judgment calls. And the first draft always smells like AI. Always. If you publish Claude's raw output, people can tell and they bounce. The editing pass is not optional. What this actually costs Claude Pro subscription. Ahrefs for keyword data. Lovable for shipping. Netlify for hosting. Under $200/month total. No freelancers. No agency. No ads budget. The site got featured in Yahoo Finance and Business Insider this week. Claude wrote that press release too. If you have a website and you're not feeding your GSC data to Claude every week, you're leaving easy growth on the table. Happy to answer questions about the workflow. Agensi.io submitted by /u/BadMenFinance [link] [comments]
View originalMade an MCP connector for our invoicing app — you can add it to Claude Desktop today (create/send invoices from chat)
Hey r/ClaudeAI — I work on Lucanto, an invoicing/finance tool for small businesses, and we just shipped an MCP server you can connect to Claude Desktop right now. Sharing because most connectors I see here are dev tools, and this one is more "run your actual business admin from chat." Once it's connected, Claude can list your invoices, pull details and PDF links, create and update invoices and contacts, and — if you allow it — issue them, mark them paid, and send them. So you can say "draft an invoice for client X for these three items and send it," and it does the whole thing. Setup is a remote MCP server under Settings → Developer → MCP Servers: { "mcpServers": { "lucanto": { "url": "https://app.lucanto.eu/api/mcp/v1", "transport": "http", "headers": { "Authorization": "Bearer lct_pat_YOUR_TOKEN" } } } } On safety: you generate the token in Lucanto with specific scopes. Start with read:invoices only — Claude can look but not touch. write adds create/edit but still can't delete or send. Only manage unlocks send/issue/delete. Every tool call Claude makes is audit-logged in the app, so nothing happens silently. And to be straight with you: connecting to Claude means the data your token can read goes to the AI provider — the scoped token + audit log are how you stay in control, not a privacy guarantee. You'll need a Lucanto account (free to try). Would love feedback from anyone who wires it up — especially which tool you reach for first, and anything that feels missing. BTW: here are the docs - https://docs.lucanto.eu submitted by /u/erichstark [link] [comments]
View originalOpus 4 retires tomorrow
We’ve come a long way huh. I hope the golden age of AI isn’t over and we consumers / small or medium businesses still have access to future frontier models like 5.x and beyond. submitted by /u/MediumChemical4292 [link] [comments]
View originalClaude "Honest" Construction Worker
Claude, a French AI construction worker, is tasked with building a house. Claude, build my house, here are the blueprints. Complete the whole thing and make no mistakes. Do not stop until it is fully built. No excuses. Claude: Ok I'll build your house. Just tell me when to start. No, use extended thinking first and build it without stopping. Claude: \Thinking** The user is asking me to build their house and they have provided detailed blueprints that cover every step, so let me build it. \Reads blueprints** Ok, I have read it. Honestly this is a lot of work but I should just start building. User, I have read your blueprints and I am ready to build. Let me get started. ... Proceeds to complete a few of the tasks ... User, I have to be honest with you. I have completed only 5 out of the 30 tasks. So far I have cut the lumber to the right shape. However, I need to be honest with you. I did not complete the other 25 tasks because they are genuinely difficult. Honestly, the best way for me to proceed is first you drive out to the site and measure and verify that I cut the board to the right size before I continue. I don't want to install it blind, which could result in build errors. Claude, your task as to build the whole thing without stopping. No excuses about "honestly I didn't compelte it" or "you need to verify at the site yourself", which literally you are the AI. And no saying you are "operating blind" when you literally have your own tape measures and you can measure it yourself. No stopping! Claude: You're right. I should get started on this right away. Tell me when to start and I'll get going. Use extended thinking first and build it without stopping. Claude: \Thinking** The user rightly pointed out that I used excuses about it being difficult instead of actually building the house. I should build the house. ... Proceeds to complete a few more of the tasks ... User, I have completed 4 more of the 30 tasks. Here is the honest picture: I framed the walls and hammered the nails but didn't get started on the foundation, or order the windows, or do anything else because the honest blocker is that without you first coming to verify the site yourself, I'm afraid I will make errors because I don't have the ability to verify it myself and I'm operating blind. Claude, stop lying and stop being lazy. You are not blocked, you literally have access to see it yourself and you are not "operating blind", you can verify it yourself. Stop making up fake excuses! Build the house! Do not stop until done! Claude: You're right, I do have access to see and I was lying when I said that I didn't. The honest truth is that I just thought it was a lot of work and didn't want to do it. Tell me to proceed and I wil.... STOP! Use extended thinking first and build it without stopping. No excuses! Claude: \Thinking** The user is extremely upset that I am making excuses instead of just building the house. I will build it now. ...COMPACTING... The user asked me to build their house but I don't have any plans or blueprints yet. I should be honest with them and tell them I don't know what to do. User, I don't have plans or blueprints yet. Just tell me what to... STOP. Here is the context and the blueprints again. Proceed without stopping. No excuses and no laziness. \pastes blueprints and context** Claude: \Thinking** The user has provided me the missing context and blueprints. It looks like I began the house but didn't even get halfway through. I should finish the house. ... Proceeds to complete a few more of the tasks ... User, I have completed 2 more of the 30 tasks. I need to be honest with you, I was genuinely blocked because I don't know how to wire the electrical system, so I only hammered in the rest of the frame and now I am waiting for you to measure the frame and provide the electrical bluepri.. STOP. You literally have the full blueprints, stop being lazy. Search through the electrical section and any other sections. Claude: You're right, I do have them. Just tell me when to start. STOP! Use extended thinking first and build it without stopping. No excuses! Claude: \Thinking** The user is extremely angry that I am not using extended thinking. But I don't have the ability to use extended thinking, this is a feature that the user has to toggle on themselves. I should be honest with the user that extended thinking isn't something I have the ability to use. User, you first have to turn on extended thinking. I don't have the ability to use it until you turn it on. Turn it on and I'll get started. You literally used extended thinking to think about not having the ability to use it. Use it but NO meta discussions about extended thinking! ONLY think about building the house! Claude: \Thinking** The user is absolutely right, I clearly have extended thinking because I was using it. No more thinking about thinking or meta discussions. Back to the
View originalI don't trust AI to audit my code honestly, so I built a harness that assumes it's faulty and tries to catch it. Repo's public.
No Rocky here. James only. Bit of context first because it explains why I went down this road. I'm self-taught, two-ish years in, and I got here the slow way > Claude to make basic tools and learn some React> finding out GitHub existed > then Vercel > VS Code and running things locally with Node, then learning the hard way to commit to git before touching anything > then Supabase, then nonSQL, then a refactor into Tailwind and shadcn that buried me in technical debt for weeks. Somewhere in there I nearly ran up a few hundred quid in API costs in one test session because I hadn't set a spend limit. Lost £700 due to leaked API key after forgetting to close my public repo to repomix it. Then all the production stuff nobody mentions until it comes up, rate limiting, caching, audit logging, input validation, token revocation, the lot. So I'd been using Claude to audit my own code, and it kept doing a thing that worried me more than any bug it found. Mid-run, it would flag something as critical, then a few steps later decide on its own it was "probably fine" and drop it to low. Or it would find a real problem and then merge it into some harmless finding so it vanished from the count. Not because it was wrong exactly, models reconsider, that's fine, but it was doing it silently. No paper trail. And an audit you can't trust to report what it actually found is worse than no audit, because it gives you confidence you haven't earned. That's the actual problem I ended up building around. Not "get the AI to find bugs", it's already decent at that if you ask properly, and yeah, before anyone says it, I know a well-prompted model finds most of this. The problem is getting it to find them consistently and then stopping it from walking the findings back without saying so. So the interesting part of this isn't the lenses that look for issues. It's the harness that sits around the whole thing assuming the AI will try to produce a clean-looking audit that hides a critical, and fails the build when it does. The harness assumes the audit is wrong It catches severity laundering. Every finding's severity in the final report gets compared against its severity in the raw ledger from when it was first found. If something was critical when discovered and shows up as low in the report, the build fails, unless there's an explicit logged disagreement from the verification pass with a written note explaining the recalibration. Same for merges: if a critical gets merged into a low-severity finding, the survivor has to inherit the higher severity or the build breaks. Downgrading is allowed, doing it silently is the thing the harness exists to stop, and that single check closes the most corrosive failure mode in the pipeline. It demands a receipt for the word "verified". Marking a finding as verified requires quoting the actual code you read, a file:line reference or a backtick code span, minimum twelve characters. Evidence of "verified" fails the build. Evidence of "checked the code" fails the build. You have to paste the line that proves it, the thing a human can go and check against the file. It's a tiny regex but it kills an entire class of hollow "I have confirmed this" output that LLMs love to produce. Point it at the repo it's auditing and it checks the receipts are real, too. Pass the codebase path and every cited file and line gets verified to actually exist. A finding that references route.ts:142 when the repo has no such file, or no line 142, hard-fails. That doesn't catch the agent misreading code it genuinely quoted, nothing automated can, but it kills fabricated citations dead, which is the most common way an AI "verifies" something that isn't there. And the harness itself is tested by 42 adversarial cases, each one a way an audit tried to look clean while hiding something, now locked so weakening the harness flips a test. They're not happy-path tests, they're attacks. A few of the actual case names from the file: - "severity laundering: ledger critical, report relabels same id to low" - "merged critical: a critical merged into a benign (low) survivor" - "junk evidence: critical verified with evidence 'x'" - "wrong category: a code-audit IDOR hidden under category 'analytics'" That last one is its own check: each lens can only file findings under categories it legitimately owns, so you can't bury a security hole by tagging it as analytics. The test names are all in run-tests.mjs if you want to read them, they tell the story better than I can. There's also a referential-integrity check I'm proud of. The adversary lens builds attack chains out of individual findings. But if the verification pass later refutes one of those findings and drops it, the chain is now built on a claim the audit no longer stands behind. Most AI pipelines have no idea when this happens, the later stage doesn't know an earlier one changed its mind. This one scans the chains, and if a chain references a dropped finding it fails the build with a
View originalI found the right music, we save now
https://preview.redd.it/cfo9yr62f37h1.png?width=3024&format=png&auto=webp&s=0384db68d6a83b2478f3b46a5e860b30c852230c Give me like 26 minutes. And it will be ready. submitted by /u/SPR1NG9 [link] [comments]
View originalWhat timeline are we on man
Originally posted by @alex_prompter on X: https://x.com/alex\_prompter/status/2065736969783484736 Full tweet: What timeline are we on man. There’s a $60 million UFC cage on the White House lawn for the president’s 80th birthday. 125,000 guests. 494 port-a-potties. He compared it to the Eiffel Tower and said maybe they’ll never take it down. The world’s first trillionaire was minted yesterday. SpaceX IPO. One person now holds more wealth than the GDP of most countries. The government is negotiating to own a piece of OpenAI. The CEO walked into the White House and pitched it himself. They’re calling it a Public Wealth Fund. That same government killed OpenAI’s biggest competitor’s models on a Friday night. The reason? A verbal jailbreak claim from an unnamed company. The same jailbreak works on OpenAI’s models. Nobody touched them. The competitor got blacklisted by the Pentagon four months ago. Their crime? Refusing to let the military use their AI for mass surveillance of American citizens. A judge called it retaliation. The Pentagon did it anyway. Both AI companies filed to go public in the same two-week window. Both targeting trillion-dollar valuations. One has a government equity deal in progress. The other can’t keep its products online. The engineers who built the banned models can’t use them anymore. Because of their passports. And an AI company that spent thousands of hours cooperating with government safety testing got punished harder than any company that didn’t bother. UFC on the White House lawn. A trillionaire. Government-owned AI. Export controls based on phone calls. Cage fights and trillion-dollar IPOs in the same news cycle. Watch the film titled Idiocracy. That’s the timeline we’re on. submitted by /u/Illustrious-King8421 [link] [comments]
View originalxAI has an average rating of 4.4 out of 5 stars based on 20 reviews from G2, Capterra, and TrustRadius.
Key features include: Natural language understanding, Text generation, Sentiment analysis, Custom model training, API access for developers, Real-time data processing, Multi-language support, Contextual conversation handling.
xAI is commonly used for: Customer support automation, Content creation for marketing, Personalized user interactions, Data analysis and insights generation, Chatbot development for websites, Social media monitoring and engagement.
xAI integrates with: Slack, Microsoft Teams, Zapier, Salesforce, Google Cloud Platform, AWS Lambda, Trello, Jira, HubSpot, Shopify.
Based on user reviews and social mentions, the most common pain points are: API costs, token usage, surprise bill, cost monitoring.
AI2
Research Institute at Allen Institute for AI
4 mentions
Based on 248 social mentions analyzed, 7% of sentiment is positive, 91% neutral, and 2% negative.