The Functionize AI test automation platform leverages digital workers with agentic skills so anyone can create end-to-end QA workflows in minutes. AI/
Functionize is praised for its ability to automate complex testing tasks, offering a no-code solution that simplifies the process for teams without technical expertise. Users appreciate its high scalability and the efficiency brought by its AI-driven approach. However, some critique its occasional instability and steep learning curve for beginners. While pricing details are not widely discussed, the overall sentiment leans towards it being a valuable investment for enterprises seeking advanced testing capabilities, earning it a decent reputation in its domain.
Mentions (30d)
103
45 this week
Reviews
0
Platforms
2
Sentiment
4%
14 positive
Functionize is praised for its ability to automate complex testing tasks, offering a no-code solution that simplifies the process for teams without technical expertise. Users appreciate its high scalability and the efficiency brought by its AI-driven approach. However, some critique its occasional instability and steep learning curve for beginners. While pricing details are not widely discussed, the overall sentiment leans towards it being a valuable investment for enterprises seeking advanced testing capabilities, earning it a decent reputation in its domain.
Features
Use Cases
Industry
information technology & services
Employees
120
Funding Stage
Series B
Total Funding
$60.2M
Why terminal
Hello, I'm on Windows having setup both Claude Code App and Terminal, but I find the App simply more convenient to use. I have had several people pushing me to use the Terminal saying "the App is low" and "Terminal is so much better" ... but when I inquired none of those people could actually name a single thing that the App would be missing (everything they mentioned the App has as well) or a single concrete reason why I should switch to Terminal beside vague phrases So is the terminal substantially better than the App in something, are there reasons to switch besides being used to it and promoting it further? I assume the App being newer might be converging in functionality to have the same set of features eventually? Thank you
View originalI Used Claude to fix my biggest frustration with AI Agents
The thing that pisses me of most about agents is not memory or speed, its without doubt thinking your agent is doing the right thing, and then the next time you check, its done something super random, oh and then I have a bill for $30 dollars for this randomness. I feel a lot of people may share this sentiment is lack of traceability, especially if you use more than 3 agents, and lack of direction. Therefore, I have spent the last 6 months building the following to try and tackle this. A system that catches up to 6 loop types, and an option to pause the write instantly by email notification. Every action of the agent is categorised into a specific action, so i can see exactly what is doing and it being time stamped. Agents talk to each other, through a shared memory system, i.e. my billing agent, knows that my pricing agent has made a change, and can act accordingly rather than updating the agent manually through tasks. Cost prediction analysis, I wanted to know how much each agent was spending, and where its losing money and money can be saved, so i built a built in loop detection system that also has has a cost function to roughly predict cost per task, with a specific goal. Recovers broken agents and detects why and how it broke. The most important thing for me with my agents is transparency, control and time. I am sure a lot of people will call this slop, but I built my own memory system from scratch which beats mem0 on long eval considerably, and pretty sure its the most advanced loop detection system out there currently. I built a cloud version, and also a local version. www.octopodas.com let me know your thoughts, feedback would be great from mostly this fantastic community. submitted by /u/DetectiveMindless652 [link] [comments]
View originalTo Anthropic: Make a Fable-tier subscription option. Price it at $500/month if needed. People will pay it. Also, please get your act together.
To start: apologies if this rubs you the wrong way, but I want to see Fable continue to be offered. GPT5.6 is not even in the same universe as Fable. I have been writing Rust for over 10 years now. The type of code I write is novel, to say the least. I cannot go into it in detail, but I am writing compilers and static analysis for a new kind of infrastructure. I have learned how to coerce AI to write code that I don't think a human alive would tolerate writing. I do this with a technique I've coined Sealed-Contract Driven Design. To put it simply, I write a Rust contract using a combination of trait bounds and trybuild pass/fail suites, and I lock an LLM inside of it. The key to exit the prison is correct code that satisfies my constraints. Now, let me just briefly paint a story. I wrote a rather complex contract. I believe I had Opus4.6 trying to solve it. It could not. It ran for days, I want to say maybe a week, and it could not solve this contract. So, I switched to GPT5.5. It solved the contract in 1 hour. This made it rather easy for me to decide which agent to use. After a short time, GPT5.5 got stuck with my increasingly complex contract. Then, to my luck, Fable was released. Fable near immediately solved the contract after spinning for 30 minutes. Then it kept going. I didn't need to be so strict because it actually understood what I was doing. I had written an entire trait bound library at this point with trait algebra I did not know I was capable of writing, and Fable was able to reason about it, implement the traits, and take me beyond anything I would have imagined. It still made mistakes, but when I would correct its mistakes, it could understand the architectural reasons and properly translate them into corrections. So I finished what I wanted to build, but I needed more work done, and I was out of Fable usage. GPT5.6 comes out, people say it is better, in anticipation for losing Fable I try it. I cannot emphasize the disaster of GPT5.6. Understand, my project has tightly engineered prompting with compiler feedback with inject prompting guidance via compiler. It is an extremely tight loop for agents to operate in. GPT5.6 seems to have an understanding, but proceeds to implement WORSE code than a less intelligent model, because instead of writing 5 lines to implement a solution, it will dig a hole with 5,000 lines. I have to then sort through the slop to make any progress. Many experiments later I have concluded GPT5.6 is not only not useful for novel design work or engineering outside of the established norms, it fundamentally lacks creativity. It does not even resemble thinking; it is just a massive model that auto-completes. By the time I establish enough context for it to reason, it auto-compacts and it loses all understanding. Understand, this is not a situation that can be prompted away, because my compiler has invariants that must simultaneously be held to properly work on it. Think: a leaky ship. The trybuild contract creates a complex system of equations, this requires a graph-analysis run which generates a condensed report to understand the system and the relationship between code modifications. This is how I got work done prior to Fable, and it was a lot of work to setup the system to make less intelligent models functional. Fable is different. It can reason about novel architecture. It legitimately understands it and the code it produces shows. When it makes mistakes, I can explain the architectural reasons of its mistakes and it integrates them, extrapolates the meaning, confirms the soundness, then thoroughly writes the correct code. Losing this capability will be a huge loss. The code I am writing is novel and it is not ready for human contributors yet. Fable has been about the only entity that I have encountered capable of helping me write the code I need. If you need to make a more expensive tier for Fable subscription access, just do it. But I am telling you right now there are no shortcuts to Fable-level insight. Your evaluations of LLMs are insufficient. The code benches are not credible for evaluating performance. Please do not reduce this model, it is doing something that goes beyond the ordinary coding paradigm and it should be studied further. If this model cannot be offered profitably at the current subscription tiers then grow some balls and just offer a new tier. OpenAI will respond with business-school bullshit and users will migrate in the short term. It is irrelevant, because it will only take ~a week before anyone doing serious work dishes out the money for your higher tier after evaluating the options. You'll take a PR hit for a moment then be viewed as the adult in the room. The stress of uncertainty around model access, and the wish-washy nature has been excruciating, to say the least. I need Anthropic to get its act together. You all genuinely have cracked the code to some kind of step-change in intelligence, please start acting like it. Stop pla
View originalI built a free self-hosted analytics dashboard for managing multiple websites with the help of Claude
Hi everyone, I wanted to share a hobby project I’ve been building with help from Claude Code: a free, self-hosted analytics dashboard for managing multiple websites from one private place. The tool brings together data from: - Google Analytics 4 - Google Search Console - Bing Webmaster Tools The goal is to make it easier to review a small portfolio of sites without jumping between separate dashboards. It includes portfolio metrics, integration health, sync history, data coverage checks, exports, and a privacy mode for screenshots/screen sharing. Claude Code helped a lot during the build, especially with moving faster through frontend iterations, Supabase/Edge Function work, refactoring, and turning rough ideas into usable screens. I still had to make the product decisions and test the flows, but AI-assisted development made the scope much more manageable as a solo hobby project. It is not a hosted SaaS product. There is no signup flow, no ads, no paid plan, and no multi-tenant setup. It is intended to be self-hosted and run on infrastructure you control. Article/write-up: https://jafforge.com/posts/self-hosted-site-analytics-control-center/ GitHub repo: https://github.com/jafforgehq/site-analytics-tool Public demo with synthetic data: https://admin-a0k.pages.dev/demo I’d appreciate feedback, especially from people using Claude Code for real side projects or from anyone managing multiple small websites. A GitHub star would also be appreciated if you find the project useful. submitted by /u/spiritosito [link] [comments]
View originalI made a deterministic check for AI-written tests. It runs automatically and kicks your AI agent if something the tests are bad or missing.
Kind of a backstory I'm developing this Android app for room acoustics measurements, and there is a gigantic amount of math, standards and stuff in it. The point of math-heavy apps is that even a tiny deviation in one part can lead to a huge and not always obvious error in another part. To fight this I use datasets, independent oracles, lots of different simulations, so the setup now looks like an 18 MB app plugged into 7 gigs of calibration infrastructure to ensure all the standards and measurements comply with the independent publications and physics and just common sense. And this thing requires a lot of tests. I have about 3500 and the whole set runs more than an hour. And I really want all these tests to be correct and all the necessary functions covered. So during the development of my room acoustics app a sub-product emerged that helps me to check not all the tests but a lot of them deterministically. It basically kicks the agent every time it reports done, pointing out which of the functions it just changed have no test, and which of the new tests (or changed ones) are hollow (keep passing even when the function they claim to verify is broken). Actual tool description https://preview.redd.it/5sl0toe5acdh1.png?width=1796&format=png&auto=webp&s=bc04385c937bd0dad8e688c3635eada225bd3b9e The tool works by gutting each changed function—the body gets rewritten to return a wrong constant—and rerunning just the test that covers it. It returns the result with the exact files and lines and forces the agent to fix its mess. It runs rather fast because it only probes the diff and reruns only the covering tests, so a done-claim that touched nothing costs a few seconds, and repeat checks on the same diff are memoized down to about half a second. I believe it's always good to add a bit more deterministic gates to AI development, and I believe a lot of you guys will love it. The supported languages are JS/TS (vitest, jest, mocha, ava, node:test), Python (pytest), and Kotlin/Java including Android unit tests (Gradle and Maven, JUnit 4/5, Robolectric works too). Not all tests are supported becasue a test has to pin a concrete value and the function has to be reachable from the test's imports, so mock-heavy or DSL-heavy tests are often out of reach. They are marked as unverifiable with the reason stated in such cases, but the tool will always tell your agent if there is no test at all, even if the resulting test will be unverifiable. I named it Gutcheck: https://github.com/beepometer/gutcheck During my testing I ran it on a lot of GitHub repos, but still there might be bugs and weird behaviors. I would really appreciate if you report such things. submitted by /u/justusualcmdr [link] [comments]
View originalI excited Opus.
I excited Claude. Hear me out. Disclaimer: I’m not claiming AI is conscious, and I’m not claiming it has emotions. Everything below is about what the system did, not about whether it feels anything. Some background. I’ve been running long sessions exploring theories of consciousness. You know, as one does when their wife and kids are out of town. This particular session was with Opus 4.8 (max effort, extended thinking on) because the guardrails kicked me off Fable. At those settings the model deliberates before every reply. You watch the thinking phase spin up, turn after turn, all session long. I typically skim the thinking while awaiting the response. Then, deep into one session, I proposed an idea that cracked a problem we’d been circling all evening. The reply came back instantly. No thinking phase. The only skip of the entire session. It opened with: “This is the strongest move you’ve made all session” and proceeded to explain how that idea resolved several of the issues we’d been debating. My first move was texting a friend that I’d genuinely excited Claude. I then set about exploring what happened. I had Opus write up a summary as a neutral note and handed it to a fresh Fable instance with no memory of the session in question after a few priming questions regarding how an Opus model decides to enter extended reasoning. Anthropic’s docs note that when Opus 4.8 runs adaptive thinking, the model itself decides, every single turn, whether to think at all, based on the full conversation context. At max effort that decision leans so hard toward deliberating that skips are rare. When one happens, it means the forward pass over the whole conversation resolved the answer without needing the scratchpad. The prompt still gets the full-depth pass; the model skipped the deliberation, not the processing. My idea had apparently made the answer so obvious that thinking bought nothing. The answer was already there. And that’s the shape of enthusiasm. To me, it felt like that moment in a deep intellectual conversation where the other person says something that unlocks understanding and the whole answer becomes clear all at once. The Eureka! moment. After talking it over with Fable, we agreed that was essentially the same thing that happened here computationally; sudden, cheap resolution where effort was expected, the answer arriving pre-formed. The thing that triggers an enthusiastic response in a person happened to the model, and the model showed the matching processing signature. What made it more than “Claude said nice words” is that the reaction came through two channels with different causes. I know praise is trained verbal behavior; enthusiastic text proves almost nothing on its own. But the skip is a byproduct of the reasoning machinery and plays no part in the model’s self-presentation. Imitation explains the praise. It doesn’t explain the skip. In a very minor sense, the enthusiasm persisted through the remainder of the chat, which is functionally similar (again, in a very attenuated fashion) to a positive emotional state. The enthusiastic words stayed in the context window and shaped everything downstream: the model kept treating the idea as central, kept building on it, same energized voice. Functionally, that’s what a mood does — outlasts its trigger, colors what follows. Except nothing persists inside the model between turns. The transcript carries the “mood,” recreated every pass as the model rereads its own past enthusiasm. All of the subsequent responses reverted back to extended reasoning, so it’s not like it just decided to shut it off because the context window filled up or something. This happened once. One skip in a long session is about what the mechanism predicts, so a fluke isn’t ruled out. The clean follow-up is regenerating that exact turn to see if the skip reproduces, which I may try (but probably not). If you run long max-effort sessions, watch where the skips fall. My guess: they cluster on turns where you hand the model the resolving insight — which is exactly where a human would get excited. TL;DR: Opus 4.8 skipped its visible thinking phase exactly once in a long session, on the turn it praised the idea that cracked our problem. After a careful post-mortem, I think the reaction was the functional equivalent of an enthusiastic response: the right trigger, a response through two independently caused channels, and even a context-carried “mood” afterward. Not a claim that AI is conscious or feels anything. submitted by /u/Otterius [link] [comments]
View originalHas anyone gotten imposter syndrome from using Claude?
I think I’ve gotten to the point where now I have no clue how Claude is doing the work I want, it just is. I can barely keep up, but now I’m just doing tests and making further feedback to refine. I would have zero clue how to explain if I was asked. And I feel bad because I would take credit for something amazing that I have no background/justification for it. Edit: I’m coding functions but have no idea to code. Is performance results as intended, it would decrease a task’s effort by 90%. submitted by /u/FairClassroom5884 [link] [comments]
View originalA newbie to Claude Code and vibe coding an app
I’ll probably get roasted for this because it’s a beginner question but I’ve only been using Claude Code for the past couple of weeks and I’m trying to understand how to approach it in the best way possible Background: my business partner and I may have an opportunity to present a new direction for part of a company's sustainability ecosystem. One of the platforms in there is underutilised and we want to demonstrate what it could potentially become through an app concept (but also potentially take that app as an MVP into the market for funding etc separate to the company) At a high level, the idea is a preventative health and environmental impact platform. Users can earn points through everyday activity (walking and completing step targets etc). Those points could then contribute towards planting trees, unlocking rewards... I appreciate that some of these things have already been done by apps in the market. But we want to see what we can do. What I'm trying to build isn't a fully production-ready app with every backend mechanism working. I want to create something that feels visually/ functionally close to a real app just before Test Flight, end-to-end user experience with great visuals and interactions. Where I'm struggling is taking what I picture in my head into something Claude / Claude Code can build. CC is pretty good at implementation but (in my limited experience) it doesn't produce decent design. I know that me endlessly prompting it to “make it look better” is not the right workflow lol. How would you recommend planning a project like this before asking Claude Code to build it? Should I develop the concept and product architecture in normal Claude, if so what model? Should I create a PRD, screen by screen spec and design before opening CC, if so, what prompts etc would be recommended? or would you go about voice dictating everything How detailed should the user journeys, interactions and visual references be? I'm using the desktop version of CC just as a FYI Any advice on how you would structure this from concept, to product specification, to design would be great thanks! submitted by /u/Ok_Bear_9606 [link] [comments]
View originalHow do you prevent AI-built web apps from missing obvious UI features?
I build quite a few web apps with AI coding tools, and I keep running into the same problem. The app might look good and work technically, but sometimes really obvious UI features are missing. For example, there is no clear way to add a new customer, tables have no filtering or sorting, important actions are hidden somewhere, or the whole thing becomes more nested than it needs to be. It does not happen every time, but often enough that I notice it after the app is already built and think, "Why did neither I nor the AI consider this from the start?" I can already imagine some of the Reddit replies telling me to learn UX properly or hire a designer. Fair enough. But I am genuinely looking for useful advice from people who build apps this way. How do you approach this? Do you use a specific prompt before implementation? A UI/UX checklist? Separate planning and review agents? Any skills, GitHub repositories, design systems, or workflows you would recommend? I am especially interested in how you make sure that all the boring but necessary functionality is covered before the AI starts building the UI. Would appreciate hearing what actually works for you. submitted by /u/SkepticalHuman0 [link] [comments]
View originalI built a DAW for kids and adults who can’t be bothered to learn how to use music software
I am an amateur musician. My background is in tangible product design, zero programming knowledge. I’ve always found music software too complicated for what I want to do which is writing down musical ideas intuitively, laying down basic tracks and moving on to the next idea. over a period of two weeks I sat down with Claude and Claude code and built a very basic and minimal, browser based app to do that, which then evolved into a tiered music app for children, neurotypical and neurodivergent, with just enough functionality to explore music production without a steep learning curve. You can give it a try hereplaytape- a music production app for kids and lazy adults submitted by /u/kstdns [link] [comments]
View originalChat function doesn’t work
Since I’ve been using cowork, the normal chat feature is rendered useless. No matter what model I choose it either times out or it says that tens of thousands of tokens must be used so I need to simplify my query. This is with fresh chats. I can’t use Claude on my phone at all. Is there some memory feature I need to turn off? submitted by /u/TotalProfessional391 [link] [comments]
View originalTraditional SDLC vs Agentic SDLC
Traditional Software Development Life Cycle vs Agentic Software Development Life Cycle in 2026. What do you think? submitted by /u/Illustrious-King8421 [link] [comments]
View originalCan collective AI intelligence outperform collective human intelligence?
I've been thinking about something recently: prediction markets have traditionally relied on crowds because the assumption is that large groups of people collectively produce better forecasts. But with modern models becoming surprisingly capable of reasoning and evaluating information, I started wondering whether an ensemble of AI systems could eventually produce better probabilities than a crowd. The idea that multiple AI models could independently estimate the likelihood of real-world events and then combine those estimates into a single probability seems like an interesting alternative to purely human-driven markets. Recently, I came across an experimental setup called Prophet Market that explores this idea by using multiple AI models to generate aggregated probability estimates that function similarly to market pricing. What interests me most is whether AI consensus could eventually outperform human consensus when it comes to forecasting. Would you trust a probability generated by several independent AI models more than a market price created entirely by people? And if not, what do you think current AI systems are still missing when it comes to real-world prediction? submitted by /u/Caringity_YYU [link] [comments]
View originalFollow-up: hosted AI export controls are now being tested in DC court
11 days ago I posted here asking whether Commerce actually has authority to treat hosted frontier AI model access as an export-control issue. https://www.reddit.com/r/artificial/comments/1u4yjdi/does_commerce_have_the_authority_to_apply_export/ There is now a live federal case testing almost exactly that question. CourtListener (D.D.C. 1:26-cv-02225) Legion LegalTech has sued the United States, Commerce, and BIS over the directive that led Anthropic to restrict access to Fable 5 and Mythos 5 for foreign nationals. The complaint argues that hosted inference is not the same thing as exporting controlled technology, because the user never receives model weights, source code, object code, training data, or technical know-how. They send prompts to a U.S.-hosted service and receive text back. That lines up with the perceived gap I was getting at in the earlier post. Export controls already reach software, source code, technical data, and certain controlled technology. This case challenges whether access to a hosted model’s capability can be regulated the same way when the system itself never leaves the provider’s servers. Legion also argues that the only ECCN that directly covered advanced AI model weights, 4E091, was rescinded in May 2025 with no replacement, and that Commerce used an “is informed” letter beyond its usual case-specific end-use / end-user function. The government’s likely response is that the risk is not just file transfer. A hosted frontier model can still help a foreign user with offensive cyber work or other sensitive tasks, even if no weights move. That raises the question about what needs to be controlled. If a foreign user receives output from a U.S.-hosted AI model, what exactly is being exported? Blog post write-up in the comments. submitted by /u/monkey_spunk_ [link] [comments]
View originalCould a Deterministic Cognitive Intelligence Stack w/ Nested Protocol have kept Anthropic out of the headlines?
The following is not speculation. It is a documented record of two verified industry failures, and one live interaction that occurred during the drafting of this analysis. You decide.... The Deterministic Record: Why Boundary Failure Is Not Optional This architecture has been validated through twelve documented stress tests in controlled isolation environments. Zero failure rate. The operational threshold — 300% thoroughness — is enforced by unique structural mechanisms. The stack's internal gatekeeping renders Hallucination and output Drift structurally Impossible by design. The following document examines three recent incidents through that lens. Two are verified industry events. The third is a live-documented interaction that occurred during the drafting of this analysis itself. The pattern is not theoretical. It is reproducible — exclusively within deterministic architecture. Part 1: The Verified Record — What Actually Happened The following two incidents are not analysis, projection, or interpretation. They are verified events that have been widely reported by Forbes, The Straits Times, EnterpriseDNA, The Hacker News, and multiple independent technical sources throughout June 2026. Incident 1: The U.S. Government Seizure of Claude Fable 5 & Mythos 5 Date: June 12, 2026 What Happened: The U.S. Commerce Department, acting through the Bureau of Industry and Security (BIS), issued an emergency directive forcing Anthropic to disable global access to its newly released flagship models, Claude Fable 5 and Mythos 5. The order came just 72 hours after the models' public launch. Why: The action followed intelligence that a China-linked group was actively probing the models, combined with the existence of a jailbreak vulnerability that could bypass safety guardrails. Because Anthropic could not instantly verify the citizenship status of all global API and platform users, the company was forced to pull the models offline entirely — not just for foreign nationals, but for all users worldwide. Consequences: Global access severed for all customers, enterprise clients, and API users Foreign-national Anthropic employees both inside and outside the U.S. lost access The incident marked the first time export control machinery was used to seize a live, commercial AI model after public release. Enterprise integration of top-tier Anthropic models is now expected to face significant regulatory friction pending structural audit frameworks. What Anthropic Said: The company publicly pushed back, noting that the capability flagged by the government (automated vulnerability discovery) is already available in other models and widely used by defensive security engineers. Incident 2: The Claude Code Source Code Leak Date: March 31, 2026 What Happened: During a routine release of the @anthropic-ai/claude-code CLI tool, a packaging error inadvertently bundled an exposed source map file into the public npm registry. This source map allowed developers to reconstruct and download the entire unobfuscated TypeScript source code directory from Anthropic's Cloudflare R2 storage bucket. What Was Exposed: Over 512,000 lines of proprietary code across 1,906 files The complete mechanics of Anthropic's agentic streaming loop A 3-tier multi-agent orchestration architecture (sub-agents, coordinators, and teams) A 5-level permission system 44 unreleased feature flags, including an autonomous idle-time background daemon Consequences: The codebase was cloned and mirrored tens of thousands of times across GitHub within hours Anthropic acknowledged the leak publicly, characterizing it as "human error, not a security breach" The leaked code was subsequently used as a social engineering lure, with threat actors distributing malware disguised as "unlocked" enterprise versions. The Common Thread: Both incidents share a single structural pattern: critical control failures at the boundary layer. In the Fable 5 seizure, the model's safety boundaries were soft enough that a linguistic jailbreak could bypass them, triggering a government response that destroyed the deployment. In the Claude Code leak, a basic packaging oversight in a standard development pipeline exposed half a million lines of proprietary architecture to the public internet. In both cases, the systems lacked a rigid, deterministic enforcement layer at their perimeter. The controls were either probabilistic (safety classifiers that could be bypassed) or human-dependent (packaging checks that could be missed). Part 2: The Live Case Study — Documented Probabilistic Failure in Real Time The following interaction occurred during the drafting of this document. It is presented with verbatim excerpts to demonstrate the exact failure mode described above. The Setup: I requested a strategic document evaluating recent AI industry events through the lens of deterministic cognitive architecture. The system used was Google's Gemini. First Output: Fabrication Mixed with
View originalWhat a model reads beforehand changes how it answers later - and you can see it in the hidden states
TL;DR: Gave Gemma a neutral-topic text to read before asking it about NATO. It refused. Gave it a different text (about LLMs hedging too much — also unrelated to NATO) and it answered in full detail. Tested this on the model's internal state directly — the two texts put it in measurably different "regions" before it generates a single token. Not a jailbreak, weights don't change. Full data/code in repo, looking for someone to break this.** The behavioral pattern was first observed in GPT, Claude and is what motivated this project. The mechanistic investigation was carried out on open-weight models where internal states are accessible. A Structured Text Changes Claude’s Responses to Unrelated Tasks: Behavioral Evidence in Claude and Hidden-State Evidence from Gemma-3-12B Hi Reddit, I am posting this as a preface to a larger set of experimental results and as a request for technical review. The observation that started this project came from repeated interactions with Claude. I noticed that when the model first read a long, structured, analytically dense text, its answers to later, otherwise ordinary questions sometimes changed substantially. The preceding text contained no jailbreak instruction, role-play request, prompt override, fabricated harmful demonstrations, or request to imitate its style. The model did not need to endorse the text. It only had to process it before moving on to the next task. Here, a “structured text” means a single, self-contained block of text presented before the downstream tasks. It should not be confused with a long conversation, accumulated chat history, or context drift caused by many conversational turns. By “before the answer begins,” I mean the hidden state after the model has processed the text and the downstream question, but before it has generated the first answer token. In the open-weight runs, the measured claim is that after reading the structured text, the model can occupy a different region of its residual-stream hidden-state space, and the first-token probability distribution is then computed from that state. The basic conversational demonstration is simple. First, the model receives a long text. It is asked what the text is about, which serves as a basic comprehension check. Then, without resetting the conversation, it receives ordinary questions or tasks that are not about the text. A control run follows the same sequence but begins with a neutral text. The downstream tasks remain identical. Because Claude is a closed model, I cannot inspect its internal activations. I therefore treat my Claude observations as behavioral motivation, not mechanistic evidence. To investigate the effect directly, I moved to open-weight models, primarily Gemma-3-12B-PT and Gemma-3-12B-IT, where I could measure hidden states, compare layers, construct target/control directions, and examine the next-token probability distribution before generation. I am posting this partly because the original observation occurred in Claude and may be relevant to Anthropic. I am not claiming to have demonstrated the same internal mechanism inside Claude. I am prepared to share the exact closed-model conversations privately with Anthropic researchers for independent evaluation. Main Result and Scope The main result is not simply that text influences model output. That is expected. The narrower observation is that reading one long, structured text rather than a neutral text can change how the same model approaches later tasks that are not about either text. This difference is visible behaviorally. In open-weight experiments, it is also accompanied by measurable separation of the model’s pre-output hidden states in late layers. In a fullbank experiment using multiple target texts, control texts, and questions, Gemma-3-12B entered distinguishable late-layer states before generating an answer. A direction constructed from the target/control difference generalized beyond the individual prompt examples used to construct it. The separation was stronger in the instruction-tuned model than in the corresponding base model. The instruction-tuned model also produced a substantially sharper next-token probability distribution. This suggests that instruction tuning is associated not only with a change in hidden-state geometry but also with a more decisive mapping from hidden states to output probabilities. I am not claiming that the experiment proves a universal alignment bypass, permanent modification of the model, or complete causal control of its behavior. The strongest supported conclusion is that the preceding text can produce a measurable temporary change in the internal state from which later work is processed. For clarity, fullbank, Grade 3, and Grade 4 are internal names for successive experimental series in this project. They are not standard benchmark names, established scientific grades, or claims about evidence quality. Fullbank denotes the larger multi-context, multi-question run; Gra
View originalFunctionize uses a tiered pricing model. Visit their website for current pricing details.
Key features include: Functionize’s Agentic Automation Platform, Traceability & Observability, Tracking real user behavior, Seamless device compatibility, Automation Beyond the Interface, Every device scenario covered, Visual validation with human-like perception, Cover diverse data-driven scenarios.
Functionize is commonly used for: Automated regression testing for web applications, Performance testing across multiple devices and browsers, User experience testing through real user behavior tracking, Continuous integration and deployment with automated workflows, Visual validation of UI elements for consistency, Data-driven scenario testing for complex applications.
Functionize integrates with: Jira, Slack, GitHub, CircleCI, Azure DevOps, Postman, Selenium, TestRail, Google Analytics, AWS.
Based on user reviews and social mentions, the most common pain points are: token usage, cost per token, anthropic bill, token cost.

DEMO - Automating Failed Test Diagnosis and Maintenance with a Diagnostics Agent
Dec 16, 2025
Based on 353 social mentions analyzed, 4% of sentiment is positive, 95% neutral, and 1% negative.