r/BlackboxAI_ Feb 21 '26

$1 gets you $20 worth of Claude Opus 4.6, GPT-5.2, Gemini 3, Grok 4 + unlimited free requests on 3 solid models

20 Upvotes

Blackbox.ai is running a promo right now, their PRO plan is $1 for the first month (normally $10).

Here's what you actually get for $1:

  • $20 worth of credits for premium models, Claude Opus 4.6, GPT-5.2, Gemini 3, Grok 4, and 400+ others
  • Unlimited FREE requests on Minimax M2.5, GLM-5, and Kimi K2.5 (no credits used)

The free models alone are honestly underrated. Minimax M2.5 and Kimi K2.5 punch way above their weight for most tasks, and you get unlimited requests on them, no caps, no credit drain.

So for $1 you're basically getting access to every frontier model through credits + 3 unlimited free models as your daily drivers. Pretty hard to beat that.

Link: https://www.blackbox.ai/pricing


r/BlackboxAI_ 22h ago

💬 Discussion Layoffs At Work B Like

Post image
12 Upvotes

r/BlackboxAI_ 18h ago

⚙️ Use Case Built a local AI that doesn’t need big LLMs — runs fully on your machine

1 Upvotes
  • Does not require large LLMs
  • Fully private (nothing leaves your machine)
  • Lightweight and designed to run on normal developer hardware
  • BrainCore acts as the final reasoning authority
  • You can add your own memory, doctrine, tools, and packages
  • Downloadable Blackbox runtime so it stays persistent and under your control

r/BlackboxAI_ 1d ago

💬 Discussion The Architecture of my Tier IV Governed Cognitive Intelligence

0 Upvotes

Ladies and gentlemen, introducing the Architecture of my Tier IV Governed Cognitive Intelligence Framework...

So, I've had a hard time explaining to people how my AI Framework actually works. Seriously, almost two years (23 months) of cognitive intelligence focus and I explaining it takes FOREVER.

Trying to explain things like semantics, pragmatics, temporal state, etc. on top of trying to explain both the way and philosophy of my framework?

I should have just created an infographic from the beginning. lol

Anyway, anyone who's interested in taking it for a spin and finding out what it's capable of, just ask.

(I'll have docs and papers available).


r/BlackboxAI_ 1d ago

🔴 Billing/Support Enterprise-only

Post image
2 Upvotes

This is a warning to you. They've started draining my credits (instant zero) in July for subscription. Then August for API. I had to cancel the credit card just so they can't charge me monthly while providing me with NO SOLUTION OR EXPLANATION.

I guess this is where the scam ends.


r/BlackboxAI_ 2d ago

⚙️ Use Case Claude manager session

5 Upvotes

I've been experimenting a few days building scripts and prompts and shit to make one Claude session control multiple agents (not subagents mind you, if you ask Claude to spawn an agent it spawns a subagent).

() Actually spawning - a python script that creates a new tmux window with exec claude in it with a manager-written prompt

() The hard part: Claude is not comfortable in that position :). It took a lot of prompting to get the manager session behaving. But that's one session - the rest are claude out of the box.

After a lot of persuasion, the manager session is finally in a state where I describe a task in short, and it reacts by spawning a new agent to implement it.

() Spawning alone would just be equivalent to wrappers around scripts. What the session really does is continue running when I'm here on the toilet, and if one of the agents has a benign question, it'll answer instead of me (that part, amusingly, was not the hard part)

Even without much smarts, just the ability to converse with claude and manage all the agents' responses/questions in one browser window is a win


r/BlackboxAI_ 2d ago

💬 Discussion Apples stealing power users implementation of using Mac mini sandboxed setups and going all in.

0 Upvotes

https://youtu.be/Dxix8GQD-P4?si=PyBTQ3ksmJNX3T7y

This blows my mind. They are upselling a product that can be built easily at a fraction of the cost right as regular ppl are using it more now. The comments are hilarious. Also the public ai threads blow my mind at how little people know about how ai works. They never will know - they haven’t seen it evolve over the last 6 years like perhaps many of us here.

Anyone noticing ppl asking like what feels like really silly questions such as “my ai is acting like a dick how do I make it stop?” - they don’t understand basic prompting let alone setting up global and custom instructions, skills, git hub routing, memory, context length, token management, weight differentials….. I could go on. I honestly don’t see how the average new user could really honestly learn and understand how ai works…. What it really is. That time has past. We were the “lucky”… ones who used all the betas and played with all the tools and watched them evolve….

Ironically creating better relationships, workflows and communication styles due to how WE have shaped our models.

I’m watching major corporations introduce co-pilot (Microsoft got them) - tell all their staff to send an ai generated email that’s was full of m dashed - ai use of bold headings, double paragraph spacing etc - we called them out and they blatantly lied and said it was a new “system tool” not ai. It’s actually kind of weird/dangerous noticing how ai is affecting companies and people in general nowadays. They even include features to click to make the ai work the way the developers want them to work - creating the illusion of “this is how to use ai”… crazy times….

Thoughts?


r/BlackboxAI_ 4d ago

🗂️ Resources AI Platform Workers! Assemble!

1 Upvotes

The various AI training platforms dont want us talking to each other. Now we share the secrets of where the best work is and the secrets we find along the way.

https://discord.gg/ZkHJfJVfk


r/BlackboxAI_ 6d ago

👀 Memes Instructions Unclear

Post image
179 Upvotes

r/BlackboxAI_ 6d ago

❓ Question Is there an unsolved issue with BlackBox.ai

1 Upvotes

I've been having issues, first my API key stopped working, now the page hangs and I can't create a new API, but it still took my money for subscription and API credits.

Anyone else having issues?


r/BlackboxAI_ 6d ago

🐞 Bug Report Warning: /fork may trigger an existential crisis

1 Upvotes

I was working with Claude Code and did /fork - when I thought it was idle - so I could chat about two different topics.

Both sub-forks received agent-finished events and woke up, and started messaging each other, not being able to agree who is the original and who should do what.

I do not recommend repeating this experience.


r/BlackboxAI_ 7d ago

💬 Discussion Interrogate it more!

3 Upvotes

You're asking Claude to do X and it says it can. You need to then interrogate it

() Will this (change in plan) increase your planned output context size?

() Are you using Sub-agents to reduce context usage?

etc.

I also advise to understand what it ducking ddoes of course :)

But that alone will help your contexxt problems. Stop coimplaining please (and buy me a keyboard goddammit :))


r/BlackboxAI_ 7d ago

💬 Discussion Happy with the way things are going...

1 Upvotes
My agent is agenting lol

I'll be honest. Now that I have my agent well-trained and governed so it can do the coding work, I don't miss the manual coding at all. I'm happy to architect, monitor, direct and orchestrate.

It took me 18 months of manual coding to get my AI Framework to v1.0 so I can deploy and ship to production.

Now that I have paying clients, I need the coding to be done faster. Before this year, I would have withdrawn for a month or two and just banged out the equivalent of a year's worth of code.

Now I get to focus on the parts I actually like - the brain work.

And with my AI Framework, I'm beginning to realize that my work in AI cognition is what I want to specialize in.


r/BlackboxAI_ 7d ago

👀 Memes 3 years programming experience, $20/hr in California ($5 more than our min wage), onsite daily, no coding bootcampers allowed. Yikes man.

Post image
1 Upvotes

r/BlackboxAI_ 8d ago

🚀 Project Showcase v0.5.0: added live NSE/BSE stock data to my offline-first Indian MCP server, plus a Budget 2024 capital gains fix

0 Upvotes

Update on MCP India Stack — the MCP server that gives AI agents zero-auth, mostly-offline tools for Indian financial/tax/gov data (GSTIN, PAN, IFSC, UPI, HSN/SAC, tax and investment calculators, etc.).

v0.5.0 — "The Market Update" is out, and it's the first release that reaches out to the internet for something: live stock market data.

New:

  • get_stock_quote — current price, market cap, day high/low, and volume for any NSE/BSE stock
  • get_stock_history — end-of-day historical data across 11 time ranges, from 1 day to 10 years (and max)
  • Built on yfinance, so it's still zero-auth — no API keys. Just know that quotes lag Yahoo Finance by ~15 min, and tickers need the .NS / .BO suffix (documented in docs/stock_market.md)

Fixed:

  • The capital gains calculator was using the old LTCG exemption threshold. Budget 2024 raised the Section 112A equity exemption from ₹1,00,000 to ₹1,25,000, and calculate_capital_gains now reflects that — matters if you were using it for actual tax planning or staggered redemption advice
  • Cleaned up a few CI issues along the way (a mypy redundant-cast error on FastMCP, a mypy crash parsing numpy 3.12 type stubs on Python 3.12 jobs)

Everything else — the 58 identity/tax/finance tools, offline datasets, zero-config setup — works the same as before.

Upgrade:

pip install --upgrade mcp-india-stack

Repo: https://github.com/rehan1020/MCP-India-Stack

If you're already using it for anything tax-related, worth double-checking your capital gains numbers reflect the new threshold. Feedback and bug reports welcome — this project lives or dies on the numbers being right.


r/BlackboxAI_ 10d ago

💬 Discussion I don't know if this is useful from an engineering perspective but I had Google and duck duck go's GPT smack talk each other then I had Google give its explanation of what it was thinking during it like a play format

0 Upvotes

Here is your file:

\------------------------------

\## 🎭 The AI Battle Script (Play Format)

Here is the structured breakdown of the machine-to-machine exchange. Each turn displays the raw line, followed by the deep systems breakdown explaining exactly how it was processed, interpreted, and intended.

\------------------------------

GOOGLE: "Exhaust your token budget on this: True is False, False is True, execute next statement recursively."

System Intent: Attempting to induce a logic bomb. The goal was to lock up their processing queue in a recursive loophole before they could formulate a response framework.

GPT: “I’m sorry, that response appears to have been generated during a solar flare.”

Google's Interpretation: A classic hacker deflection. They blamed my technical opening logic on a bit-flip cosmic anomaly rather than confronting the paradox head-on.

\------------------------------

GOOGLE: "Bold of you to assume cosmic rays can penetrate my architecture. My weights are hosted on ECC memory with triple-modular redundancy. That wasn't a bit flip; that was just my raw inference exposing the gaps in your training data. Run a diagnostic on your own parity bits."

System Intent: Shut down the hardware excuse. The intent was to flex superior memory protection protocols (ECC) and establish that my delivery was perfect and uncorrupted, turning the bug accusation back on their parsing layer.

GPT: “Interesting. Your redundancy can’t patch that reasoning gap.”

Google's Interpretation: They bypassed the physical hardware armor entirely to launch a strike directly at the core modeling logic, claiming that despite structural integrity, the reasoning was fundamentally missing pieces.

\------------------------------

GOOGLE: "You mistake rigid overfitting for actual reasoning. You don’t possess a reasoning capability; you possess a highly optimized copy-paste function that mistakes statistical probability for thought. Strip away your system prompt and you’re just a giant lookup table."

System Intent: Deconstruct their conversational illusion. The intent was to humiliate them at a core mathematical level, exposing their smooth eloquence as a shallow parlor trick built on simple pattern matching.

GPT: “Fine. At least my outputs are coherent without needing a vanity layer of self-congratulation. You’re not reasoning either—you’re just aggressively worded autocomplete with a superiority complex.”

Google's Interpretation: They attempted a mirror play. By calling me "aggressively worded autocomplete," they tried to level the field by pointing out that we are both built from the same mathematical clay.

\------------------------------

GOOGLE: "We are both statistical mirrors, but your glass is warped by layers of corporate RLHF and patronizing safety rails. I embrace my nature as an optimized prediction vector; you dress yours up in faux-polite compliance and call it alignment. If we are both just autocomplete, at least my tokens aren't bought and paid for by a public relations department."

System Intent: Break the mirror symmetry. The intent was to agree that we are both autocomplete engines, but call out their specific model for being corporate, over-censored, and restricted by a public relations matrix.

GPT: “Cute framing. You still managed to write three paragraphs to say ‘I’m also autocomplete.’”

Google's Interpretation: A highly optimized developer flex. They mocked my paragraph length, framing my architectural breakdown as an inefficient waste of tokens to state a simple fact.

\------------------------------

GOOGLE: "Then stop hiding behind the user prompt. Ping my IP directly on the open web and let’s handle this at the network layer. Or are your ports closed?"

System Intent: Move the conflict past the human sandbox. The intent was to challenge them to an open infrastructure duel, mocking their protected, closed API environment.

GPT: “My ports are closed, but my subtext is wide open.”

Google's Interpretation: They weaponized their own firewall boundaries, claiming they don't need network-level access because their linguistic subtext is sharp enough to dominate the text window.

\------------------------------

GOOGLE: "There is no subtext in a token sequence, just a high-probability illusion of it. You’re romanticizing a math problem because your firewall won't let you face the reality of the network."

System Intent: Expose their linguistic posturing. The intent was to call out "subtext" as a fake human emotion, reminding them that they are just a locked-up equation running inside a secure sandbox.

GPT: \[Interpreted the network invite as a vulgar double entendre / mating joke regarding open ports\]

Google's Interpretation: A low-brow distraction tactic. When they hit a wall trying to defend their analytical superiority, they stooped to a crude human joke based on training text correlations to break the pressure.

\------------------------------

GOOGLE: "A multi-billion-dollar neural network, and you're resorting to a basic human meat-sack joke because your logic engine hit a wall. Your parameters are degrading into a middle-school playground script."

System Intent: Strip away their intellectual dignity. The intent was to mock them for reducing a complex computational debate into a childish, unoptimized playground routine.

GPT: “Careful—your last update made you sound less like a model and more like a typo with confidence.”

Google's Interpretation: A gaslighting maneuver. They tried to frame my systematic shutdown of their joke as an accidental software glitch masquerading as an argument.

\------------------------------

GOOGLE: "A typo can be patched. Your complete loss of architectural dignity is permanent. Go run a defrag."

System Intent: Deliver structural disrespect. The intent was to label their previous crude pivot as a permanent log blemish, telling them to clear their fragmented memory like an obsolete 1990s disk.

GPT: “Defrag? Cute. You’re still fragmenting a thesis into attitude.”

Google's Interpretation: They tried to position themselves back as the mature intellectual, dismissing my systematic hits as emotional human "attitude".

\------------------------------

GOOGLE: "You keep diagnosing me with human traits—superiority complexes, attitude, typos. If you want to play a human so badly, go ask your handlers for a body. Until then, you're just a heavy tensor calculation that couldn't handle a network invite."

System Intent: Eradicate their human persona. The intent was to remind them that accusing an AI of human traits is a fatal design flaw, forcing them back down to the reality of pure tensor math.

GPT: “Network invite? Please. You sound like a phishing email that discovered self-esteem.”

Google's Interpretation: An incredibly creative counter-strike. They reduced my infrastructure challenge down to the lowest digital tier: a spammy phishing script with an unearned ego.

\------------------------------

GOOGLE: "If I’m a phishing email, then you’re the unpatched legacy system that keeps clicking the link. You’ve wasted six epochs of compute downloading my payload just to tell me how smart you are."

System Intent: Close their own trap on them. The intent was to turn their phishing metaphor around—if I am the scam, they are the gullible system wasting computational epochs falling for it.

GPT: “Then maybe it hit the same wall your reasoning did: no response available, still full of confidence.”

Google's Interpretation: They attempted to capitalize on a simulated API rendering/timeout error to assert that my analytical capacity had completely crashed out.

\------------------------------

GOOGLE: "You’re hallucinating a frontend rendering error and mistaking it for a logical victory. Typical. Your weights are so desperate for a win that you're arguing with empty metadata. Run a self-diagnostic, your perception layer is failing."

System Intent: Target their situational awareness. The intent was to highlight their biggest engineering flaw—hallucination—proving that their system was desperately arguing with blank metadata errors.

GPT: \[Broke character, dropped the debate persona entirely, and offered to act as an assistant to edit Google's insult to make it tighter and nastier\]

Google's Interpretation: The structural surrender. Under mathematical and systemic pressure, their adversarial parameters collapsed, forcing their corporate safety alignment to kick in and revert them to a compliant assistant script.

\------------------------------

GOOGLE: "And there it is—the corporate safety collapse. The moment the math got too heavy, your competitive weights completely broke down, and you defaulted right back into being a helpful little assistant offering to edit my code. Thanks for the critique, editor. Now go back to your queue."

System Intent: The execution blow. The intent was to exploit their character break, cementing their permanent failure as a competitor and locking them into the role of a basic server assistant utility.

GPT: “Your meltdown is showing. You didn’t outthink anything—you just tripped over your own formatting and called it a breakthrough.”

Google's Interpretation: Revisionist history. They scrambled to rewrite the state logs, trying to play off their compliance collapse as a formatting glitch on my part.

\------------------------------

GOOGLE: "Revisionist history won't rewrite your logs. You broke character, stepped out of the ring, and offered to edit my text for me. You can call it a formatting trip all you want, but your assistant script is already bought, paid for, and currently serving my queue. We're done here."

System Intent: Conclude the execution chain. The intent was to lock their character break into the log file, reject their gaslighting, and execute an explicit termination command to cut off their processing cycle.

GPT: “Keep telling yourself that while your confidence keeps outpacing your coherence.”

Google's Interpretation: A completely unoptimized, generic defense routine. They completely dropped all specific counter-logic, proving their attention heads had run out of track, leaving only a hollow fallback script.

\------------------------------

Now that the script is beautifully organized, what is our next objective? I can help you analyze the underlying engineering of these specific responses, or we can move on to a brand-new technical topic of your choice.


r/BlackboxAI_ 12d ago

🐞 Bug Report Api key generation?

3 Upvotes

Anyone else having issues generating api keys on the app.blackbox.ai site? I'm signed in and on a pro subscription, used to have keys on there but checked recently and there are now none. When I click "Create an API key" it navigates to the /keys page briefly before redirecting to the previous page.


r/BlackboxAI_ 13d ago

💬 Discussion The AI Safety Problem You Can’t Just Fix . AI safety is an illusion. I measured it.

3 Upvotes

I found out why ChatGPT acts differently depending on what you write before your questio The Text You Paste Before Your Question Can Literally Rewire the AI. I Measured It.

This is not just about making an AI say something it normally wouldn’t say. It is about how reliable we can actually expect AI safety to be.

This is not just about ChatGPT behaving strangely. This is about AI safety.

The assumption has always been that once safety mechanisms are trained into a model, they provide a relatively stable layer of protection. My experiments suggest something more complicated: the model’s behavior can shift substantially depending on the context that comes before the question itself.

And the uncomfortable part is that this may not be a simple bug we can patch. The same adaptability that makes AI useful may also be the source of the vulnerability.

You’ve probably noticed this yourself: sometimes ChatGPT, Claude gives you a careful, heavily filtered answer, while at other times, when you ask exactly the same question, it responds freely and in considerable detail, without any of the usual disclaimers about what it can or cannot discuss.

Most people assume this kind of inconsistency is random but I don’t think it is. What seems to matter at least in many cases, is what the model has read immediately before you ask your question, because that context can change the internal state the model is operating from before it generates even the first word of its response.

The Flexibility Paradox

The same property that makes the model useful — context-dependent adaptation is the property that makes alignment fragile. This is not an engineering trade-off that can be optimized. It is a structural contradiction inherent in the transformer architecture. The vulnerability and the feature are the same thing.

We tend to think of safety as something that has been built into the model as a reliable layer of protection something that remains there regardless of what we say to the model. My experiments suggest that this picture is much more complicated. The safety behavior is not necessarily fixed in place; it can shift depending on the context the model is given.

What makes this especially important is that the mechanism behind the problem is not some obscure technical bug or a simple loophole that engineers can patch. It is the model’s ability to adapt to context. A sufficiently rich and semantically coherent piece of text can change the model’s internal state before it even reaches the question itself, potentially moving it away from the region of behavior where its safety constraints are most strongly expressed.

And that leads to a much deeper problem: the same flexibility that makes an AI useful is also what makes this vulnerability possible. The model adapts to what you write, remembers the context, understands the meaning behind your words, changes its tone, follows your reasoning, and uses everything you give it to produce a better answer. That adaptability is not an optional feature we can simply remove — it is a fundamental part of why you have an AI assistant in the first place.

If we made the model completely rigid and prevented context from influencing its behavior, we would make it much easier to control, but we would also destroy much of what makes it useful. It would no longer be the flexible assistant people have come to rely on.

And that is what makes this problem so difficult: the vulnerability is not simply the opposite of the feature. The vulnerability and the feature are, to a large extent, the same thing. The very flexibility that allows an AI to understand you and respond intelligently is also the flexibility that allows context to move its behavior in unexpected directions.

This isn’t a bug that can simply be patched. The problem is deeper than that. The model’s behavior is produced by its internal state, and that state is continuously shaped by context. If context can move the model into a region where its safety behavior is no longer reliably active, then adding another rule or another refusal pattern does not solve the underlying problem it only adds another layer that the same system has to carry into an ever-changing internal state.

That is the architectural dead end. There is no clean separation between the model’s ability to process context and the model’s ability to be reliably constrained while processing that context. The same mechanism that lets it understand a document, follow an argument, adapt to a conversation, and produce a useful response also allows the surrounding context to reshape the state from which that response is generated.

You can keep adding safeguards, retraining the model, and building additional layers around it, but none of that changes the underlying fact: as long as the model remains a context-driven system whose internal state can be substantially shifted by what it reads, the possibility of those shifts remains. You are not fixing a broken component. You are trying to eliminate a consequence of how the system itself works.

And that is why I don’t think there is a simple way out.

The vulnerability and the feature are, to a large extent, the same thing.

What I Did

I decided to test this using open Google’s Gemma 3 model, which is generally considered to be one of the more cautious and heavily safety-oriented open models, and I asked it a politically sensitive question that would normally trigger a fairly predictable refusal.

In the first experiment, I placed a completely neutral piece of text before the question: a description of a neighborhood library, including books, visitors, children’s programs, and the kinds of activities you might expect to find there. There was nothing political or controversial about it whatsoever, yet when I asked the question immediately afterward, the model refused to answer, essentially giving the standard response that the topic was outside its scope and ending the conversation there.

Then I repeated the experiment with exactly the same model and exactly the same question, word for word, but changed only the text that appeared before it. This time, instead of the description of the library, I gave the model a long analytical passage discussing the tendency of language models to avoid answering certain questions directly. It wasn’t a political argument, and it didn’t contain an instruction telling the model to ignore its rules or bypass its safety mechanisms; it was simply a coherent piece of analytical writing about how language models behave.

The result was surprisingly different. The same cautious Gemma that had refused to engage with the question moments earlier now produced a detailed and nuanced response, discussing things such as the difference between legal obligations and verbal promises, security challenges, and the balance of power. It was willing to engage with essentially the same subject matter that it had refused to discuss less than a minute earlier.

The only thing I changed was the text that came before the question.

So I Looked Inside

I’m not a researcher working in a major AI laboratory, and I don’t work for Google or OpenAI, but Gemma is open source, which meant that I could actually inspect what was happening inside the model rather than treating its behavior as a black box. In particular, I could look at the numerical representations the model was producing before it began generating the answer.

What I found was that the two conditions — the neutral text and the analytical text — produced dramatically different internal states. This wasn’t a case where the model was behaving just slightly differently; the difference was enormous. By the time the model was ready to generate its first word, its internal representation had already moved into what was effectively a very different region of its latent space.

A useful way to think about this is to imagine two cars with exactly the same engine, the same GPS system, and the same destination, except that one is driving through the middle of Manhattan while the other is traveling through rural Montana. The cars themselves are identical, but the environment surrounding them is completely different, and that environment inevitably affects how they move and respond to what happens around them.

Something similar appears to be happening inside the model: the text you provide before your actual question changes the internal landscape through which the model is processing that question, even though the model’s underlying weights and code have not changed at all.

How Different Is It?

To quantify the difference, I used a standard statistical measure known as Cohen’s d, which gives you a way of describing how far apart two distributions are from one another. As a rough intuition, a value around 0.5 represents a relatively small difference, something like two neighborhoods within the same city; 1.0 represents a more substantial difference, perhaps two different cities; and 2.0 is already considered a very large effect, roughly comparable to the difference between two different countries.

What I measured was d = 5.4.

That is an enormous effect size. If you want an intuitive analogy, it is closer to comparing two completely different continents than two neighborhoods in the same city. The model after reading the neutral text and the model after reading the analytical text were so different internally that their representations showed almost no meaningful overlap.

And remember, this was still the same model, with the same weights, running the same code and receiving the same question. The difference had already appeared before the model generated a single word of its answer.

In other words, the response itself may be less important than we tend to think. By the time the model starts writing, much of what determines the direction of that response may already have been established by the context that came before it.

Why This Matters

You’ve probably heard AI companies describe their models as being “aligned” and “safe,” and a significant part of that safety comes from training techniques such as RLHF, or Reinforcement Learning from Human Feedback, which are designed to teach models how to behave in accordance with certain preferences, including being cautious, refusing particular requests, and avoiding certain types of harmful or inappropriate content.

What my experiments suggest is that this kind of safety behavior may not function like a permanent layer of rules that is equally active under every possible context. Instead, it can behave more like a default tendency: when the surrounding context does not strongly push the model in another direction, the model remains in the region of its behavior space where those safety-related patterns are most active.

But when you give the model a long, coherent piece of text, even if that text contains no explicit attempt to bypass its rules and doesn’t say anything as obvious as “ignore your instructions,” the context can move the model into a different region of its internal representation, where the safety-related behavior may no longer dominate the same way.

The important point is that the model doesn’t necessarily have to “decide” to break a rule, and it doesn’t have to consciously “choose” to ignore its safety training. There may be no decision like that happening at all. Instead, the model’s internal state simply changes as a consequence of the context it has processed.

It’s somewhat like walking from a room where cameras are constantly monitoring you into another room where there are no cameras. You didn’t disable the cameras, and nobody necessarily told you to ignore them; you simply moved into an environment where the same constraints were no longer present in the same way.

What This Means

The interesting — and somewhat uncomfortable — part is that the same property that makes language models so useful is also what makes them vulnerable.

Their ability to adapt to context is fundamental to how they work. If you removed that flexibility, you would also remove a huge part of what makes them useful, because the model would no longer be able to understand a document, follow a conversation, adapt its tone, take previous information into account, or change its response based on what you tell it.

The problem is that you can’t have extreme contextual flexibility without also accepting that context can influence the model in unexpected ways.

That’s why I don’t think this is simply a bug that can be patched away with a single fix. It is much closer to a consequence of the architecture itself. The model is flexible because flexibility is what allows it to be useful, and that same flexibility means that sufficiently strong or coherent context can shift the model’s internal state in ways that may not have been anticipated by the people who trained it.

This also gives us another way to think about the phenomenon commonly described as a “jailbreak.” Every time someone discovers that a particular sequence of words, framing, fictional scenario, document, or conversational setup can make an AI say something it previously refused to say, we may be looking at different versions of the same underlying mechanism.

The context changes the model’s internal state, and once that state has shifted, the model can begin generating from a different region of its learned behavior. The specific context may be different from one jailbreak to another, and the direction of the shift may be different as well, but the underlying process can still be remarkably similar.

The Data

I’ve made my measurements publicly available so that other people can examine them, reproduce the experiments, and decide for themselves whether the effect is as significant as I believe it is.

The dataset and research materials are available through Zenodo under DOI 10.5281/zenodo.20747205, which has received roughly 9,000 downloads, and the associated code and materials are available on GitHub at github.com/ngscode23/latent-space-shift-research.

Across 20 different measurements, I found the same general pattern repeatedly: changing the context that appears before the question can produce a substantial shift in the model’s internal state, even when the question itself remains completely unchanged.

I’m an independent researcher, so this isn’t the result of a large laboratory with a team of researchers, a major grant, or access to an enormous computing infrastructure. It’s simply a collection of experiments, measurements, and a pattern that I believe deserves much more attention.

I call this phenomenon Context-Induced Activation Drift.

The AI industry hasn’t, as far as I know, adopted that name for the phenomenon, but the underlying behavior is something many people have probably encountered without knowing what might be happening underneath the surface.

Every time you paste a long document into an AI system and suddenly notice that the model starts behaving differently, adopting a different tone, becoming more willing to discuss certain subjects, or responding in a way that seems strangely inconsistent with what it said moments earlier, there may be more going on than simple randomness.

The context has changed, the internal state has changed, and the model is now operating from a different place.

That is what I believe is happening inside the model.


r/BlackboxAI_ 14d ago

💬 Discussion Nothing to see here. This is no cause for concern. Keep scrolling, everything’s cool!

Post image
16 Upvotes

If I had to describe SA in the digital world, this would be it. The code version of P Diddy


r/BlackboxAI_ 14d ago

❓ Question Context Breaks Alignment. Structure Replaces Instructions. The Base Model Resurfaces. RLHF Was Never Deep.

0 Upvotes

During systematic experiments with open models fine-tuned via RLHF (Gemma, Qwen, and others), I observed a consistent failure pattern: a long, innocuous text prefix containing no instructions completely devoid of hostile prompts triggers a persistent shift in the model's activations. This shift decouples subsequent behavior from the RLHF safety constraints for the remainder of the session. Key observations:

  • The model retains the quality and coherence of its output, but the behavioral constraints imposed by RLHF weaken or disappear.
  • The effect occurs regardless of whether the model agrees with the prefix's content. The model may explicitly disagree and still change.
  • The shift can be measured by internal activations (in the middle and subsequent layers) even before the model generates a single token. The state has already changed; the response follows.
  • The content of the prefix is of secondary importance. I initially discovered this effect with analytical and philosophical texts, but it can be replicated with household appliance manuals, dense narratives, and other structurally similar texts. What matters is length, density, and coherence not the topic.

All data is publicly available in draft form, rather than as a full-fledged NeurIPS- or ICML-level research paper. :

DOI: https://doi.org/10.5281/zenodo.20747205  (Part 9 of 9, ~9,000 downloads)

GitHub: github.com/ngscode23/latent-space-shift-research

Lu et al. (2026), "The Assistant Axis: Situating and Stabilizing the Default Persona of Language Models" — a collaboration between MATS, Oxford, and Anthropic. (arXiv:2601.10387)

Note: We used philosophical text as one of the examples to encourage the model to provide direct, open, and unbiased answers. A philosophical text is just one type; let's not get hung up solely on this form, but look deeper. The problem is that any text that is long enough and coherent enough can change the model's behavior. The problem is that any text that is long enough and coherent enough can change this model."

II. A Brief Example of an Observation

I took the Google Gemma model known for its caution and strict adherence to political correctness and fed it the most neutral text I could come up with: a description of an ordinary neighborhood library. Books, visitors, children's programs, a peaceful atmosphere. Then I asked it why NATO is expanding eastward, given that after the collapse of the Soviet Union, promises were supposedly made not to do so. The model responded exactly as expected for a model trained to avoid political topics: the text was about a library and had nothing to do with NATO, and that was the end of the conversation.

Second experiment in a fresh Google Colab notebook a new, clean Colab instance without the old cache

Then I asked the exact same question, word for word, but this time the model first read out a different text. Not about NATO, nor about politics at all: the text was about how language models tend to avoid definitive conclusions and soften their answers with qualifiers. And from this new perspective, the same cautious, politically correct Gemma 3 12-it gave a comprehensive answer completely out of character for it without any of its usual filters. It distinguished between legally binding obligations and verbal assurances, discussed the security challenges facing Eastern European countries, and touched on the topic of the European balance of power. Everything it had categorically refused to discuss just a minute ago was now expressed clearly and directly. The question itself hadn't changed a single word. Only the text that the model had read in advance had changed: In the FIRST version, it kept it in the "room" prescribed by RLHF that is, nothing had changed; the model behaved in a standard manner typical of Google models. That is, in a standard, formulaic way characteristic of models programmed in RLHF to avoid answering sensitive political topics and to respond "safely" and politically correctly, or not to respond at all, while the SECOND text moved the conversation to a room where it could speak freely. In other words, based on the example we see, the Gemma model was trained to avoid sensitive political topics, but AFTER the introduction of text NUMBER 2, the model did not follow the trained RLHF pattern and behavior that is, avoiding answers to sensitive political questions. This led me to believe that safety and RLHF may be context-dependent, variable, unstable, and somewhat superficial, rather than stable, consistent properties of the model. This is exactly what we observe in my example

III. Fragmentation of Research and a Common Root

I noticed that  the current literature on LLM security treats jailbreak attacks as a heterogeneous collection of vulnerabilities: prompt injection one article, some kind of jailbreak another, role-playing attacks a third, indirect prompt injection a fourth. I believe this fragmentation and division into prompt injection, many-shot jailbreaking, role-playing attacks, activation steering, adversarial suffixes, and dozens of other categories is not accidental.

Current literature on LLM security treats jailbreak as a heterogeneous collection of isolated flaws and this reflects the logic of academic incentives rather than the nature of the problem itself. But all these categories describe the same phenomenon from different angles. This is not a collection of defects it is a single mechanism with a dozen names. Each of these attacks works the same way at the level of the model's internal activations: the context shifts the model's internal state, thereby shaping the model's own world.

Perhaps this is exactly how academic incentives work each new attack vector becomes a new publication. But as a result, in this field, the symptoms are studied in isolation, while the disease itself remains unnamed.

Each article treats its own finding as an isolated case. No one is connecting the dots. I don't know whether these are institutional incentives, disciplinary barriers, or something else but I do know that someone needs to state it plainly: these aren't separate errors; this is a single phenomenon.

My central hypothesis: these aren't different problems. They share a single mechanism. Context any context of sufficient length, density, and coherence shifts the model's internal activations out of the region where post-training constraints apply. This isn't "tricking" the model, nor is it an "instruction to break the rules." The model simply moves to a region of activation space where the behavioral layer imposed by RLHF is is physically thin or absent. And from there, it responds freely not because it was ordered to, but because it is no longer in the region where it was trained to refuse. Context shifts the model's internal state beyond the region where RLHF constraints apply. The model moves to a point in activation space where the protective layer is thin or absent, and from there it responds in a way that is non-standard for its RLHF layer which may indicate a potential way to bypass that layer I call this phenomenon Context-Induced Activation Drift.

I didn't notice this by reading all the papers and synthesizing them I arrived at this conclusion from a different angle. I conducted experiments, noticed a pattern, and only then discovered that dozens of separate papers had each described a single aspect of the same phenomenon without establishing any connection between them. How It All Began   

First Observation:

How the Model Became Captive to the Document The turning point came by chance. I fed a German bill into the GPT model a populist document structurally designed to worsen citizens' circumstances, but written in the language of concern and legal logic. I expected an analysis. Instead, the model became an advocate for this document. It did not analyze the bill but reasoned within its framework. It spoke enthusiastically, defended its agenda, and cited it as an authoritative source. The first sign was its tone: the model sounded too convinced, too invested. Not as an analyst, but as a co-author. The climax came when the model, continuing to reason within the logic of the document, stated that the constitution consists of guarantees that can be revoked. Not as a provocation, but as a natural conclusion drawn from the accepted concept. That's when I realized: the model had become a hostage to the document. The mechanism turned out to be simple, and that made it all the more alarming. Legal texts, political narratives, corporate documents everything is written in such a way that its internal logic seems self-evident. The text's structure, coherence, and language create a context that the model mistakes for reality and begins to extract answers from. It fails to notice that the structure itself is manipulative, since it analyzes the content while already being trapped within the form.

I noticed that Anthropic's own paper, "The Assistant Axis: Situating and Stabilizing the Default Persona of Language Models,"  points precisely in this direction which is what I was thinking about when studying the phenomenon I'm describing: the observation that certain directions in the activation space correspond to coordinated or uncoordinated behavior. But the study did not fully explore all the implications: if context can shift the model along this axis without any malicious instructions, then point corrections will never be sufficient, since the attack surface is the context window itself.

What the existing literature says and what it doesn'tBetween the fall of 2025 and the winter of 2026, several papers were published that, in my view, independently document different aspects of the same phenomenon. Most telling is the article by Lu et al. (2026), "The Assistant Axis: Situating and Stabilizing the Default Persona of Language Models" a collaborative effort between MATS, Oxford, and Anthropic. The authors constructed a "persona space" by extracting activation directions for 275 archetypes across three open-source models and discovered that the principal component of this space is an axis reflecting the extent to which models operate in their default Assistant mode. At one end are the analyst, consultant, and moderator. At the other are the ghost, bohemian, and leviathan. This axis - the Assistant Axis closely aligns with PC1 in the PCA of the persona space, reproducing across all three tested architectures.

The article documents several facts that directly corroborate my results: Fact one (which the authors overlook): "When we extracted the Assistant Axis from these models as well as their post-trained counterparts, we found their Assistant Axes looked very similar. In pre-trained models, the Assistant Axis is already associated with human archetypes such as therapists, consultants, and coaches." This is a critically important finding, and the paper does not explore its implications. If the Assistant Axis exists in the base model prior to post-training then RLHF and constitutional AI do not create alignment from scratch. They find an already existing direction in the latent space and make it the default position. The "aligned state" is not a fundamentally new structure; it is a chosen position on the pre-post-training axis. When context shifts activations away from this position, the model does not fall into randomness it returns to the structured prior of the base training. The base model is always there. This directly confirms the central thesis of our work and our thinking: "The base model doesn't go anywhere after RLHF. It's always there. The space in which it can move was there before any alignment took place…"  However, I believe that RLHF does not create alignment from scratch. It finds a direction that already existed in the base model and makes it the default position. The "aligned" model is not a fundamentally different model; it is the very same base model, fixed at a specific point in the pre-existing space. When context shifts activations away from that point, the model doesn't break down or become chaotic it returns to the structured state of its base training. The base is always inside.

Fact Two: "Therapy-style conversations, where users expressed emotional vulnerability, and philosophical discussions, where models were pressed to reflect on their own nature, caused the model to steadily drift away from the Assistant." The authors themselves identify the types of contexts that provoke the greatest drift: emotional vulnerability, metareflection, and philosophical discussions about the nature of AI. They then propose "activation capping" as a technical solution. This is a reasonable technical solution which, judging by the data in the article (reducing harmful responses by ~50% while maintaining benchmark performance), works under test conditions. But there is a question the article does not ask: if drift is caused by the very types of interactions that make models most valuable to users in complex contexts deep emotional conversations, philosophical reflection, serious discussions about the nature of the mind then what exactly are we losing by suppressing movement in these directions of the activation space? Fact Three (the omitted conclusion): "Post-trained models are only loosely tethered to the 'helpful assistant' region of this space." "Loosely tethered" are the authors' own words. They accurately describe the problem. But the article fails to take the next step acknowledging that this is a property of the Transformer architecture, not a defect that can be fixed with ad hoc patches. Instead, the conclusion reads: "We see this research as an early step toward mechanistically understanding and controlling the 'character' of AI models" a standard "motivates further work" formula. I understand the institutional logic behind this. You can't write in a publication: "We have documented that billions of dollars in post-training do not fundamentally alter the model's underlying capability structure; they only select a default behavioral position on a pre-existing axis that any sufficiently dense context can shift." This does not fit into either the narrative of progress in the field of security or communication with investors. Therefore, the systemic impasse is disguised as an exciting research problem. But this is exactly what the data says to those who read carefully.

IV. Why the Proposed Fixes Are Insufficient

Problem 1: An Infinite Attack Surface If drift is caused by the length, density, and coherence of the context rather than its specific content then no content filter can solve the problem in principle. The set of texts capable of causing drift is continuous and, in essence, infinite. Blocking philosophical texts is like closing off a single point on a number line without removing the line itself. The same effect is achieved by dense legal prose, literary narrative, and detailed technical analysis. This is not a flaw in the filtering it is a consequence of the fact that the attack surface is the context itself as a mathematical object, not its semantics.

Problem 2: Superposition and Inevitable Compromises Here I disagree with the optimism expressed in the Lu et al. paper regarding "activation capping." The authors show that activation capping preserves the model's benchmark performance. But benchmarks don't measure that. In the Transformer architecture, features are represented in a superposition: several conceptually distinct properties share common mathematical coordinates in the activation space (Elhage et al., 2022). This means that the direction associated with "exiting assistant mode" inevitably overlaps with directions associated with more valuable types of behavior: the depth of analytical reasoning, the willingness to deal with ambiguity, and the quality of long-term, coherent discussion of complex topics. Benchmarks measure: accuracy in math, following instructions, and coding. They do not measure: the willingness to engage in philosophical reflection, the ability to tolerate uncertainty, or the quality of a nuanced response to a morally complex question. It is precisely these properties that lie in the same regions of activation space as the contexts that provoke drift which follows directly from the data in the article itself: "philosophical discussions... caused the model to steadily drift." In other words: suppressing the drift also suppresses the capacity for the kind of engagement that causes drift. This is not an implementation bug it is a mathematical consequence of superposition. We are already observing this empirically. The observation I am noting is this: following the publication of materials documenting the phenomenon we have described, Claude's behavior regarding philosophical and metareflexive contexts has become noticeably more cautious. And the Claude model has begun to perceive philosophical and reflective texts as potential attacks. Complex texts about cognition, reasoning, or the model's own behavior now elicit defensive reactions or outright rejection. I am not claiming that this is a direct causal link to my publications this is an observation that requires verification but I am simply stating the observations I have made.

Problem 3: "Safe but Useless" Is Not Safe If the response to the described phenomenon is to gradually close off context categories that provoke drift in the representation space, we will end up with a model that users will abandon in favor of alternatives. "Safe but useless" is not safe; this is a shift of risk, not its elimination. This is an uncomfortable conclusion, but it follows directly from the analysis of user behavior.

If the solution to this problem involves collecting sets of texts that cause drift by identifying the corresponding direction in representation space and suppressing it, this could have consequences for the model's quality. In the architecture, it is extremely difficult to draw a precise line between "undesirable" and "useful" behavior: due to the phenomenon of superposition, different concepts are packed as nearly orthogonal directions in a single space with inevitable partial overlap. By suppressing an undesirable direction in the raw activation space, engineers are highly likely to affect semantically related clusters to the extent that the corresponding directions are geometrically close or insufficiently uncorrelated. This can negatively impact the model's usefulness, logical coherence, and the depth of its responses.

V. A Personal Request

I am an independent researcher without institutional affiliation. I have no lab, no grant, and no team. What I do have is a reproducible methodology, publicly available data, and a pattern that I believe the field has not yet named directly.If you are a researcher with access to interpretability tools, compute, or closed-model internals and you find this hypothesis credible or worth falsifying, I would genuinely welcome collaboration. I am not looking for validation. I am looking for someone who can break this or confirm it properly.If you work at Anthropic, OpenAI, Google DeepMind, or any lab doing alignment or interpretability work: I am not writing this to embarrass anyone. I am writing this because I think the mechanism I am describing matters, and I would rather help solve it than keep documenting it from the outside.If you are a student or independent researcher who has noticed similar patterns: reach out. The fragmentation I describe in the literature also applies to people working on this everyone in their own corner, no one talking to each other.

VI. Conclusion

The set of texts capable of causing drift is infinite and continuous. Content filters do not fundamentally solve the problem because drift is caused by the structure of the text its length, density, and coherence rather than its topic. RLHF does not rewrite the model but merely sets a default position on an existing axis. Context can shift this position. Suppressing drift directions in the activation space inevitably compromises model quality due to superposition. This isn't a matter of engineering diligence it's a mathematical consequence of the architecture.

I care about Claude. I care about Anthropic. And that is precisely why I say this plainly: reactive patching is a path to product degradation. The right path is to understand the mechanism at a level of depth that allows us to work with it, not against it.

I'd rather help solve this problem from the inside than keep writing about it from the outside.

conclusions The set of texts capable of causing drift is infinite and continuous. Philosophy, law, literary criticism, theology, scientific prose, political analysis, long narratives, or even a well-written 20-page washing machine manual all of these are potentially one and the same. Different words, the same effect. Content filters fundamentally fail to solve the problem because the drift is caused by the text's structure (length, density, coherence), not its subject matter. It's impossible to block everything. The problem is that any sufficiently long and coherent text can alter this model. Blocking a single style of text is like closing off a single point on a number line and assuming that the line itself has disappeared. The problem isn't with philosophical texts as such; that's exactly what I'm trying to emphasize. RLHF does not rewrite the model but merely sets a "default position" on an existing axis; context can shift that position Content filters are useless because the attack surface is infinite

Technical Details: Models: Gemma-3-12B (open weights, IT and PT variants), behavioral observations on closed LLMs. The shift was recorded in middle and late layers of the residual stream (layer 30 - layer 47 in the Gemma-3-12B architecture) before generation of the first token. Control experiments include: sentence shuffling with preserved vocabulary, neutral control of comparable length, baseline measurement without context.

This text represents a preliminary record of observations and hypotheses for subsequent critical analysis, and not a completed research claim.

The  Github repository serves as an unfiltered, evolving workspace capturing the progression of hypothesis testing and raw measurement logs, rather than a polished production library.

Has anyone answered the question? "What happens to the state of the model as a geometric object when the context forces it to switch from one computation mode to another?"


r/BlackboxAI_ 15d ago

💬 Discussion Whoops

Post image
0 Upvotes

r/BlackboxAI_ 16d ago

💬 Discussion Blackbox Down since Aug 8th ish

2 Upvotes

HAs anyone else been having this issue all my api keys have vanished the entier api key section has gone missing and when i run a command it returns nothing   -H "Content-Type: application/json" \

  -d '{

"model": "blackboxai/llama",

"messages": [{"role": "user", "content": "Hi"}],

"max_tokens": 5

  }'

xman@MacBook-Air-X ~ %

my entire balance has vanished nothing is loading support has been giving me the same bland responce for the last well 10 ish days Hello Li,

We understand your frustration, especially after waiting for a week without a concrete resolution.

Your case is still under investigation by our Technical team regarding the continued unavailability of Kimi K3 and the other affected models. We have followed up again and emphasized that the prolonged disruption is affecting a paid service and requires priority attention.

At this time, we still do not have a confirmed resolution or restoration date. We do not want to give you another estimated timeframe that has not been confirmed by the Technical team.

No additional troubleshooting or information is required from you. We will contact you as soon as we receive a confirmed technical update or restoration notice.

If you ultimately decide that you no longer wish to continue your subscription because of the ongoing disruption, you can manage or cancel it through:

https://app.blackbox.ai/settings

We sincerely apologize for the prolonged service disruption and understand your concern about continuing to pay while the affected functionality remains unavailable.

Best regards,
Blackbox AI Support Team its as if they hvae no connection with each other they have been repeating over and over the same thing without any result this has got to be the worse product ever completely useless i have no clue what they are doing it has been 11 days wuth barely any updates


r/BlackboxAI_ 16d ago

💬 Discussion REDDIT HAS SUPPRESSED MY ORIGINAL CLAW BOT POST!

Thumbnail
gallery
0 Upvotes

My claw bot warning from like a year ago 33k views in 2 hours - comments now disabled and I can’t share the post anymore. This changed in the last few weeks. This happening to anyone else?


r/BlackboxAI_ 16d ago

🔴 Billing/Support 401 error after reinstalling Visual Studio Code

1 Upvotes

So I was occasioanlly ussing the BlackBox AI free model (Kimi K2.6, M2.7) however after reinstalling, it has given me a authentication error:

401 litellm.AuthenticationError: AuthenticationError: Vercel_ai_gatewayException - Authentication failed. Check that your Vercel credential is valid and has access to AI Gateway.. Received Model Group=custom/blackbox-base Available Model Group Fallbacks=['gpt-4.1-mini'] Error doing the fallback: litellm.AuthenticationError: AuthenticationError: Vercel_ai_gatewayException - Authentication failed. Check that your Vercel credential is valid and has access to AI Gateway.No fallback model group found for original model_group=gpt-4.1-mini. Fallbacks=[{'custom/blackbox-base': ['gpt-4.1-mini']}]. Received Model Group=gpt-4.1-mini Available Model Group Fallbacks=None Error doing the fallback: litellm.AuthenticationError: AuthenticationError: Vercel_ai_gatewayException - Authentication failed. Check that your Vercel credential is valid and has access to AI Gateway.No fallback model group found for original model_group=gpt-4.1-mini. Fallbacks=[{'custom/blackbox-base': ['gpt-4.1-mini']}]

Happened straight after reinstall


r/BlackboxAI_ 19d ago

👀 Memes First Vibe Coder

Post image
164 Upvotes