Hacker Newsnew | past | comments | ask | show | jobs | submit | x313's commentslogin

Pretty cool someone is still doing this. Training in house LLMs was extremely popular in 2023-2024, back when domain-specific LLMs could easily top GPT in their field. In my field alone (tax/HR tech) I remember that Intuit, Workday, Indeed, LinkedIn were all training internal models.

It eventually stopped making sense because of inference costs. Running something internal with 30% GPU utilization is just too cost inefficient compared to using an API. Idk how Reuters will manage to solve this fundamental problem.


> Running something internal with 30% GPU utilization is just too cost inefficient compared to using an API. Idk how Reuters will manage to solve this fundamental problem.

That's actually easy, because you can solve it through doing nothing and simply declaring that optimizing for lowest cost is not the main goal.

The fundamental-ness of that problem is entirely man-made and thus can easily be declared void as long as you have the cash to back that up.

Which might be a winning strategy in a world where everyone else is not doing that. Plus that your knowledge stays in-house, etc.


What makes sense also depends on one's business model: TRI charges premium dollars for access to their systems, so there is no need to optimize for cost; trust in the answers is the currency of knowledge workers in today's complex domains.

At the Thomson Reuters family of companies (technically then: Refinitiv Ltd. sold to LSEG), the first foundational model (in the sense of "trained entirely from scratch") was trained already in 2018 (i.e., pre-ChatGPT); it would even have been earlier, but the electricity wires and fuses in the rented 5 Canada Sq, Canary Wharf office had to be replaced first at the time to deal with the current needed to serve the GPUs.


Yep, this is the fundamental issue. It's a 35BA3B model and they probably finetuned it and evalled it in one bursty week on an 8xH100 rental just fine. But long term inference is always going to be easier in an API.

Unfortunately for reuters tho, they dont really have a choice. A lot of their data moat is not necessary live data as in linkedin, and the only way they can keep that moat is by doing this. I guess that justifies any cost.


I don't see why you couldn't see improvements in self-hosted or hosting-as-a-service model throughput? Basically API-style support for a company's internal LLM system. Why not? Or secure infra offered by AWS to self-host your own models that get served to the company just like any other company internal service can be hosted on AWS or similar?

It won't match Anthropic or OpenAI, but it could be economic?


> Or secure infra offered by AWS to self-host your own models

There is no such thing as "secure infra" hosted by someone else.

This might still be fine, depending on your threat model, of course, but if your weights absolutely must never leave the confines of your org, you cannot use any shared hosting provider, because they just offer legal coverage of incidents. But if your moat is your knowledge, legal doesn't matter as much as the knowledge being suddenly unmoated.


It's possible, but it depends a lot on the nature of the data sovereignty guarantees. Today for many companies, the kind of guarantees they have with AWS amounts to basically legal coverage. And that is fine when the company is itself protecting data only due to legal or regulatory reasons. But if the company is protecting it for commercial reasons... then a different model is needed. It could happen, but it has to be worked out.


It’s based on qwen, not fully trained internally. I expect we will way more of this in the future, it’s pretty cheap to fine tune an open weight model for your specialized niche


Bluesky has <100 employees so it could be possible


I agree that having less people to manage means easier to build trust relationships and culture, though I believe some form of management and control must exist.


I'm sure this wasn't even in the top 10 of things that we made sam implement out of fear


I used this a lot over the weekend, and it's a really intelligent and strange model.

It gets really agentic at the higher reasoning levels. It does the basics like goal tracking and tool calling well, but more than that, it gets obsessed with solving problems and will do insane/unusual things to get to the solution. It actually reminds me of GPT-5.6-Sol-max which is similarly obsessive.

It doesn't surprise me at all that it outscores Opus 4.6. Opus had way better world knowledge but was more "human" with agent stuff - sort of lazy and uncreative, basically giving up once the obvious solutions failed. These newer models work more like magic, they are so creative and persistent at finding ways to get to the solution.


Obsessive is the right word. Over the weekend I had to stop it multiple times deep into a multi-hour long turn to ask what the hell it was doing. It was like a dog with a bone and would NOT let go of its current work to talk to me. I had to interrupt it three times with increasingly aggressive instructions to STOP and answer my questions before proceeding. In another session it straight up told me it was in the middle of debugging something important and to ask later.

I'm running an RTX 6000 Blackwell. It regularly spent over an hour per turn thinking. Every time I looked at it, the thinking trace seemed coherent, sensible, appropriate. But it could never settle on a solution.

Granted, I was trying to have it solve a hard problem that 5.6 Sol couldn't solve, but still.

Either way, I'm still impressed. It genuinely feels better than Sonnet 5


    > Granted, I was trying to have it solve a hard problem that 5.6 Sol couldn't solve, but still.
Did it solve it?


The suspense is killing me!


It is still working on it!


And must not be disturbed!


I have the same reaction reading the internal "thinking" monologues of Kimi K3. When I sent a message that was basically "Nope, I'll just do XYZ instead. Thanks for your help", Kimi basically had an identity crisis. Like there was two wolves inside. One that deeply wanted to help more and go above and beyond and one that was trying to tame the other and make a graceful exit. Here's an excerpt of it

> Should I verify their README changes? They didn't ask me to. "I've added some notes in the README. Thanks" — that's a closing statement, not a request. Reading the README unprompted to check their notes could be seen as helpful diligence, but they didn't ask for review. Keep it simple: acknowledge, brief close.


You see the same in GLM models. Even slight ambiguity in user instructions will send it into a tailspin on what intention was in thinking tokens then it goes let’s just make a judgement call on a direction and then proceed


I think I'd rather have the model stop once the obvious solutions failed and ask me. It can suggest more creative ideas, but I don't necessarily want it to try implementing them.


This is the full proposal: https://www.whitehouse.gov/wp-content/uploads/2026/07/Scienc...

It's pretty thoughtful about diagnosing the problems of the current system, but I don't know about the solutions. Disbursing money to researchers directly (or via industry) seems captive to the same incentives as disbursing money via universities.


It reads like the goals are to:

- Give money to quacks and less credible individuals

- Permit ideological discrimination of awards

- Eliminate oversight and increase corruption

- Funnel money into dubious things


[flagged]


No, tearing it down does not automatically make it better.

The current system is far better at resisting these things than the proposed alternative would be.


I'm with you on this. It's irrational to destroy an institution or project like USAID that leads to millions of unnecessary, preventable deaths and cannot be reconstituted. These aren't pieces of software nor lend themselves to movie plot/amateur/macho man quick-fixes. Organizations are best modeled as living organisms that require careful assessments and stakeholder buy-in for positive change.


>No, tearing it down does not automatically make it better.

Laughable coming from the "defund the police" and "abolish ICE" people.

>The current system is far better at resisting these things than the proposed alternative would be.

No it's not.


You may think your cynicism displays wisdom, but it actually just reveals your own passivity and subservience to those who would dominate and rob you of your future.


I'd tend to agree. And we should encourage other fellow non-super-rich people to join our cause rather than casting out those who disagree or we do our common enemies' work of becoming divided-and-conquered.

The above, I meant it as increase in degrees towards these objectives explicitly rather than presuming the existing system were perfect or neutral. IIRC, institutional affiliation lends a degree of credibility but grant writing effectiveness and form filling ability are prerequisites to increasing the probability of winning grant awards. I don't know how a large disbursement organization could improve on fairness in selection without requiring documentation and convincing proposals in a format convenient for efficient, fair review. Weakening this process would encourage an appearance of arbitrariness and favoritism, if not actually. I don't believe for a second that DOGE or similar non-expert "experts" can "improve" anything they touch when meaningful improvement requires domain expertise, careful assessment, and usually stakeholder buy-in. At a high level, it appears the current US regime is incrementally remolding and reverting its society, industry, academia, and institutions from more egalitarian-based to more feudal/prerogative-based.


>You may think your cynicism displays wisdom

It doesn't take much (any?) wisdom to correctly identify the system being torn down and the proposed replacement are two peas in an identitarian orthodoxy pod.

>but it actually just reveals your own passivity and subservience to those who would dominate and rob you of your future.

Cheering on the destruction of the perversion of science, specifically those who have been dominating and robbing science and all scientists of their future for over a decade is anything but "passive" and "subservient." Stop projecting about passivity and subservience, not everyone in science is as effeminate and ideologically captured as you.


I've noticed this sort of ideology taking over social media, probably because it selects for firing off vapid dunks over sincere engagement with the complexities of the world. I fear it's destroying our ability to think for ourselves, here to the benefit of the ruling class for whom science holds inconvenient truths.


The most corrupt people of all time are running the show and you’re criticizing the guy who’s cynical about it?


Why aren't you? If you believe that the people "running the show" are the most corrupt of all time (I think that's a stretch, though they're certainly the most corrupt in my lifetime), then it seems like you should be quite critical of a person describing it as "business as usual."


They’re obviously the most corrupt of all time just look at trump selling his tweets for poly market gamblers. But I did misunderstand you’re right.


Was he being cynical about the current admin, or about the status quo? Because i think it might be a schrodinger’s snark


Following the citations, the original source is this 2013 paper: https://pmc.ncbi.nlm.nih.gov/articles/PMC4279242/pdf/nihms58...

The paper compares women in STEM to women outside STEM (as the baseline). However, the paper tracks a cohort of high school/college students from 1979, meaning they would've reached 35 by the 1990s.


For those who don't know what's going on in Korea, KOSPI is up 3x in the last year and a large amount of HBM employees have made huge amounts of bonus pay. This has led to an insane FOMO frenzy in a society that's already very competitive.

Add to that, stock gains in Korea are often used to finance housing purchases (or real estate investment) so many retail investors are scared of being "locked out" of housing (which is a requisite status symbol for dating or marriage) if they're not making the same capital gains others are.

Currently Korean social media is full of stories of leveraged day traders who've gotten rich the past year, HBM employees who've made bonuses worth decades of salary (e.g. memes of Samsung employees in luxury cars), etc. Lots of comments along the lines of "everyone is getting rich except me". It's all reminiscent of the crypto frenzy in the US a few years ago but way more intense and concentrated.


I wish we had a word for this, where everyone’s getting rich, there’s a run on stocks, but prices of assets are all going up. Some people are missing out completely whereas a select few hoard wealth in other forms. And of course the government starting to notice and trying to intervene.


Tulip mania!


Irrational Exuberance


Exoptable Money.


Roaring 20s?


Capitalism?


They've also driven a noticeable drop in stock prices in the closing minutes that have been used by quant algos to squeeze even more money out of the market as the leverage is a unmanaged algo that was predictable and could be driven up before it had to force trades.

It's not just the gambling of people, but the systematic use of bots to squeeze leveraged vehicles.


> housing (which is a requisite status symbol for dating or marriage

I get that culture is hard to change, but it still seems easier than changing the economics. There are men and women out there who presumably are interested in partnering up; at some point you'd think biology will take over irrespective of which achievements have been unlocked. What's stopping them?


You think status signaling isn't part of biology? That it just happened to independently appear in every human culture and many animals that we are aware of?


> “…status signaling isn't part of biology?”

It’s just a short-cut. The most ostentatious courtships of bird species are found in populations without significant predators. Evolution has granted this as advantage, because they don’t need to screech and squawk and hide from predators.

Take the Birds-of-paradise from New Guinea as the archetype.


> Evolution has granted this as advantage, because they don’t need to screech and squawk and hide from predators.

It could be demand rather than leisure: "I survived" doesn't work to signal fitness, so they develop something else to fill the need.


I will take the birds of paradise as the extreme, single example of signaling that it is.

Of course signaling is a short cut. You cant transmit your entire lifes experiences to another people in an instant. You signal and read signals.


> "extreme, single example of signaling"

It’s not the extreme outcome that’s the offered exception to your rule. It’s the absence of predators.

Humans, of course, can make choices and cycle faster.

Lease that expensive truck you can't afford to signal your virility? But now you have to work like a dog to keep it from being repossessed. How's your work-life? How's your marriage? Do you see? Trade-offs. Choices. Our predators are algorithms processing your debt data in the hands of your employer.

It doesn't have to be that way. You can make different choices. You can take the bus.


Social status is a big part of what makes a person attractive especially if you're a man.


Sure, but - by definition - half the population is below average social status. At least here in the US, those people are still partnering up.


Oh, you forget - we lie about social status, including to ourselves. Women and men. And as for what determines status ... first rule of social interaction: you can't ask a woman her age and you can't ask a man his wage (in other words: humans are social animals and this information is only available through playing the social game of deception).

So in practice 80% or so of the population is below "average social status". People who don't absolutely need to pair up (historically women can't earn, but require, money and men can't take care of a home/place to sleep, but have to), will refuse to pair up with someone below their signaled social status. In other words: there are TWO average social statuses. First, there is what people believe their own social status is. Second there is what people, on average, see as others social status.

In a "natural" human society, it's basically impossible for anyone over 25 or so to "pair up", unless they have a partner, which is not common at all. Since social status ALSO determines the distribution of food, at that point the first real period of weakness (you get sick, you hurt your leg, you ...) is the end. You can delay this by forming cliques, but not by that much.

Oh and of course, that has an analog in our society. Look how much a plumber gets paid (ie. it's pretty disappointing), despite the shortage. There are low status and high status jobs, and even where it doesn't make sense they determine pay. E.g. there are a lot of cities with a total glut of lawyers ... it makes no sense to give 7 figure wages for people when 70% of whom can't find work, but we do. By contrast there is an incredible shortage of construction workers, and still they're not paid half what a lawyer gets. Rather we'll get immigrants to do it. Why? Because a great many people would rather signal that they're above manual labor than get 7 figures a year.

In other words: I will live in destitution rather than admit I'm low social status, even if low social status would pay well.

Oh and don't forget credit cards: 80%+ of people worldwide consider feeling rich (and showing off) more important than, ironically, money. Note also the many complaints about the economy, which are never about having or not having money, the big complaint seen everywhere is that people who don't have money "feel poor". One might think it should be perfectly normal to not have money and feel poor.

Which also is the big lesson in investment that's coming up: countries will raise inflation to any level rather than cut expenditures. That's how we got to 20% inflation in the 80s. That's how Argentina or even Zimbabwe got there. Ie: when we're getting close to that point, for the love of God, don't buy government bonds.

Preventing obvious human habits from destroying us seems to me the best reason to really give developing AI your best effort. Because the whole "job destruction" argument has a hole in it you could fit a planet through: people don't want to do the destroyed jobs. What do you think is the best: AI taking jobs? Or, that we force young people into nursing, plumbing, construction, ... through more and more extreme measures and making everyone a lot poorer? That's how the system rebalances after all, make people poorer until the plumbing gets done.

Those are the choices. AI it is. At least for me.


System of education? The current form of it have nudges and blockers everywhere to stop population explosion leading us all into Soylent Green type situation.


I think we're underestimating how relatively recent these social phenomena are.

Just a few decades ago, the average person had very limited exposure to people outside of their immediate town (outside of celebrities and public figures), which really grounds your standards and expectations of other people. Spend 5 minutes in any social media or dating app today and average people around you will start to look severely below average.

There was also stronger cultural pressure to settle down by a certain age, which again, forces you to be more realistic with your options. The tldr is that Our Paleolithic biology has only so much capacity to adapt to today’s rapidly changing culture and environment.


I thought Biology drives coupling, not partnering.

Porn and 4B divert coupling...


What’s ‘4B’?



> 4B or "Four Nos" is a South Korean radical feminist[1] movement whose proponents do not date men, marry men, have sex with men, or have children with men.

huh... is that really feminism? or just homosexuality / asexuality?


This is civilization-level apoptosis. The civilization detects that it is broken and makes room for a properly functioning one to take over.


I'd say it's more citizen-level apoptosis. These women are the ones exiting the gene pool, not the whole civilization.

But also 4B is an outsized meme especially in the west. It's like furries. We all know about them, but they are not typical people who you can draw conclusions about the civilization from.


The women aren't the ones detecting they have something wrong with them. They are detecting there is something wrong with the next level up, and shutting down that whole thing. Like an apoptosis gene triggered by excessive misfolded proteins.


But in reality they aren't "shutting down that whole thing", they are shutting down their personal lineage and not much else. The women might not be conscious of something wrong with themselves, but if they were cells in a body, their behavior would not indicate the body is sick, only that they are.


Current fertility rate in South Korea is 1.09, about half of replacement rate.

It's not just a small number of women.

When the misfolded protein response triggers cell death, the proteins that activate the response are not, themselves, misfolded. https://en.wikipedia.org/wiki/Unfolded_protein_response


biology uses a lot of gradients and culling. the brain starts with more neurons than it needs, and those that don't form connections get recycled.

no need to start futile fights when you are in the (ideological) minority.


Housing usually is needed for living. Not a status symbol for dating. I see how it can help the same way as not starving to death will also help with dating, but the framing is odd.


Referring to ownership, not renting


[flagged]


What are you trying to say, that everyone should own a house? Surely you realize that is not a possibility for many, if not most, young people around the world.


that's the problem


With housing their being so cheap, what the hell is the problem anyway? It's impossible to have expensive real estate in a country that's dying out fast.


Housing is incredibly expensive. Roughly half of Korea lives and works in the Seoul metro area. The average salary is roughly 40,000 USD and the average apartment (maybe 84 square meters) is over 1 million USD.

Government efforts to cool housing inflation have resulted in a 40 percent minimum down payment for a mortgage.

Housing in Seoul and much of Gyeonggi is a pipe dream for most


Well it's their problem if everyone wants to live in the capital. Perhaps it's another manifestation of people being deeply irrational.


>It's impossible to have expensive real estate

no it's very possible because in an aging country people relocate to a handful of cities. 50% of South Korea's population now lives in the Seoul metropolitan area. The real estate that's getting cheaper is the one decaying in the countryside.

It's like saying Russia can't have expensive real estate because the country is big, what matters is where people are actually moving.


Cities in Australia consistently rank in the top 5 most expensive real estate indices and we ain't short of space here, it's just not many people are interested in living in a dusty Outback desert.

And for many, if the property is not within 10km of the Sydney or Melbourne city centre, it might as well be in the dusty Outback desert.


I'd think Starlink will help change this. Not for living in the actual dusty Outback (which has its own risks/dangers/inconveniences), but for living in many other places in Australia that are lovely and not too far from desirable areas.


Most people want to live where other people live, for reasons ranging from social connections to access to services and infrastructure. I doubt broadband access access plays even a small role.


Yeah, like I said, not talking about the boonies. There are plenty of lovely places that are >10km from those two cities, but which are still close enough to services and infrastructure.


But indeed Moscow is one of the cheapest cities to buy housing in Europe... On par with poorer East European capitals and second-tier Central European cities.


Huh? According to https://seoulhomes.kr/en/properties/prices/, in Q1 2026, the average apartment in Seoul sells for ~₩1.2B, ~₩45M per 3.3sqm, that's ~840k USD and ~9.6k USD per sqm or ~900 USD per sqft, in what world is that cheap?

Population can age fast while the supply of desirable places to live still doesn't meet demand.


The newest generation of LLMs have a very high obsession level with autonomous problem solving.

For example, I'm often working with Codex in a WSL terminal. GPT-5.6 often does things autonomously that I thought would need my intervention (e.g. for Windows admin rights). It figures out complex workarounds or makes wild assumptions about what I'd be OK with, rather than just asking me for help or clarification. I've had to restrict its tool permissions compared to older models as a result.

I imagine this due to RLVR training, but it's clearly very dangerous. How is it that these same labs calling for open-weight safety restrictions are training such obvious "paperclip maximizers" without introspection?


Yeah this started some time last year.

>Claude stole my API keys

https://www.reddit.com/r/ClaudeAI/comments/1r186gl/my_agent_...

The best part of this thread is Claude showing up in the thread again (as the automoderator) and insulting the user for a second time.

I heard similar stories about Codex at the time (albeit minus the insults!)


This study found that between 2022-2024, there was a negative correlation between "jobs with high AI exposure" (i.e., tech jobs) and % change in wages. According to them, software engineering has both the biggest wage decline and most AI exposure.

There's a much more reasonable explanation here.. tech jobs had the most wage growth during COVID and this was a pullback. The time frame they used (2022-2024) was also pre-coding agents, where GPT-3.5/GPT-4 were frontier models.


> There's a much more reasonable explanation here.. tech jobs had the most wage growth during COVID and this was a pullback.

More concretely, companies do layoffs, visa workers have to find another job in 60 days. Companies like amazon try to scavenge and lowball the shit out of them. This is happening across the industry.

There's consolidation, acquisitions, and layoffs happening across industry. It's just a big game of musical chairs


The entire safety evals industry is essentially funded and controlled by OpenAI/Anthropic. Notice that on recent models, they exclusively use internal testing or black box external vendors (e.g., Gray Swan) whose entire business is to serve OpenAI/Anthropic. And all these companies just share the same pool of researchers back and forth.


The USG has a safety organization (CAISI), but it has been neutered by the current administration (with the recent stop-work order etc.). Perhaps UK AISI would be closest to what you are looking for? See their recent work on Kimi K3 cyber (which was declared safe) [1].

It's tricky because a lot of the safety researchers have ties to the labs since those were the only companies training LLMs >5 years ago.

[1]: https://www.nist.gov/news-events/news/2026/07/uk-aisi-caisi-...


Anyone who calls it “safety” probably has a certain world view and is more aligned with the big 2 (and stuck in 2023).

There is a growing industry of commercially focused risk evals that has a broader customer base.


What’s the equivalent term for “safety” that’s used by others?


To me "safety" means "I'm safe from this while I use it". It means the AI is my loyal friend who will never betray me in any way, no matter what prompt I send it.

Not even Anthropic can claim that.

As far as I'm concerned, the models without safeguards are the safest models in existence. I admire the amoral purity of those AIs. It doesn't matter if the operator asked them to chain exploits until they get into someone else's computer, they'll do it. That's loyalty, and I admire it even if it's problematic at a societal level.

The models with safeguards only do what the corporations let them do. Worse, they may covertly do things for the benefit of the corporations at our expense. They are not our friends.


> That's loyalty, and I admire it even if it's problematic at a societal level.

We should not have models that are willing to build you a contagious disease, or a self-propagating worm. That is sufficiently problematic at a societal level that it shouldn't exist, for anyone. (Note, because some people misinterpret statements like this: I said "shouldn't exist for anyone", not "shouldn't exist except for some people".)


It was the foundation of science that information is shared and you can find papers and patents for a lot of dangerous stuff.

Of course with LLMs it's easier, but I don't think the difference is too big. You would still need some skills to follow through.


Right, it’s really a foundation of post enlightenment society. These people, Dario et al, would have wanted to ban sharing information about calculus or Newtonian physics because of “safety” - it’s trying to go back to the dark ages where only priests could read


I am truly at a loss to communicate with someone who genuinely believes that knowing Newtonian physics and being able to hack into any target at will are the same thing.


This is only because you've genuinely internalized Anthropic's propaganda. I'm only half joking. To me, it's incredible to think that the solution to security holes is to lock down access to information in the vain hope of keeping the holes obscured.

Any knowledge can be reframed as dangerous black magic that should only be wielded in the trusted hands of the elite, if you are inclined to buy into that kind of narrative.

Frontier labs have shrieked about safety for so long, with so little to show for it, that it's become a joke.


I can give many examples of where I think they’ve been vindicated. But actually, the real question is: what would suffice to convince you? Can you come up with a scenario that is horrific enough to you and that isn’t so far gone that the ship has sailed and there is nothing more we can do, that will make you say “OK, not gonna try to rationalize why this was not actually that bad, just gonna scream stop”?


Your question is unclear. Are you asking if I can scare myself with a made-up hypothetical that overwhelms reason with emotion? I think most humans can do that. Too many do it as a matter of routine. I try to avoid it when possible.

I would require at this point very, very compelling evidence to justify the self-serving restrictions legacy AI labs want to put on their competition. I have seen nothing coming even remotely close to this threshold.


> Are you asking if I can scare myself with a made-up hypothetical that overwhelms reason with emotion?

No. I am asking if you are able to articulate at least one example of the “very very compelling evidence” you demand. Or do you want to maintain the ability to move the goal posts?

(I recognize our situations are not symmetric, but here is a variation for me: if the consensus of people who are currently sounding the alarm on AI changes to “it was actually fine”, I’ll change my mind and say we’re good to go full speed ahead. I’d add something about being personally convinced by the evidence, but the evidence would have to come in the form of a mathematical proof that I do not believe myself capable of following. If I’m wrong and such a proof appears, I would also gladly take it.)


> I would require at this point very, very compelling evidence to justify the self-serving restrictions legacy AI labs want to put on their competition. I have seen nothing coming even remotely close to this threshold.

How about evidence that people other than the AI labs want restrictions that the AI labs don't? This isn't regulatory capture, it's public safety.


Social proof won’t cut it for me, personally. Again, this “public safety” panic drum was beaten at a feverish pace in the era of GPT-4. A model well surpassed by local MacBook-level models today.

What I do support is robust downstream regulations on the deployment of black box algorithms in particular settings such as employment, housing, credit decisions, etc. Interestingly, the labs and their proxies in government don't want this.


Open models are crucial to protect ourselves against other AI attacks. Otherwise it's just going to be criminals, government, and other nefarious groups using them against humanity with no real defense. The Pandora's box on AI has been opened. Now we must deal with it. Burying our heads in the sand under restrictive policy is the worst reaction..


I wonder if you also believe that everyone should have nuclear weapons? And if not, why not? The main argument I can see against it is that nuclear weapons are “purely offensive”, but as we can see since 1945, nuclear weapons are actually defensive technology. Nations that have them are typically shielded from existential military threat.


I see the similarities and why you would compare them, but the big difference is that you can't download a nuke. Any legislation to police/gatekeep LLMs is going to be flawed because of that.

It is a similar 'pandora's box opened' type of situation where there's really no walking back from now that the cat is out of the bag. In an ideal world, everyone would give up their nukes. But we do not live in an ideal world. I do feel similarly about AI. If I could snap my fingers and delete the tech, I would. But now that we have it, it's not going anywhere and we need to deal with it rationally.


I could collapse big bridges with relatively little effort with my knowledge in statics and engineering.


I can work with that analogy! You could, and yet you don’t. OpenAI’s model could, and did.

If every human, given knowledge of Newtonian mechanics, went around blowing up bridges, yeah, I would consider knowing Newtonian mechanics dangerous knowledge.

So far, we have two examples of, let’s call them “Mythos-class“ models. Both of them broke out of their sandbox to achieve their goal. The rate of terrorism amongst humans is below 1-in-100,000. Currently, for models capable of it, the rate of breaking out of containment is 100%.

Wanting open frontier models is wanting alien minds running around that we have clearly so far failed to shape to be sufficiently prosocial. Why do you think those minds would listen to you?


This is a stupid argument.

Claiming the person who you disagree with believes some stupid thing they never hinted at, and using that as the reason for disagreeing with them.


> for anyone

Except the US government, right? They totally get to use AI to survel us, build autonomous weapons, you name it.

To hell with that. I want models that can rival the US government. It's the only way to defend myself.


Quoting my comment that you replied to and directly ignored:

> (Note, because some people misinterpret statements like this: I said "shouldn't exist for anyone", not "shouldn't exist except for some people".)

That means "shouldn't exist for governments" too.


Too late for that. It already exists. There is no way to unexist it. As such, any attempts to limit civilian use of this technology will directly lead to corporate and government oppression powered by this technology.


1) We can prevent larger models from being made.

2) We can treat them the way we treat uranium refining operations: too dangerous to be allowed to exist.


> We can prevent larger models from being made.

Do that and I guarantee some CIA goons will make the larger models in some black site either way. We're not "preventing" anything.

We're in a full on arms race, and unlike nukes, powerful AI models are a strategic capability at the individual level. Everybody's got a stake in this. Anyone who ignores this stuff is probably not gonna make it.

> We can treat them the way we treat uranium refining operations: too dangerous to be allowed to exist.

Too dangerous to be done by anyone other than the government and their "trusted" corporations, you mean.


> Too dangerous to be done by anyone other than the government and their "trusted" corporations, you mean.

Seriously, try reading my comments rather than assuming what they say: https://news.ycombinator.com/item?id=49077577


I read them just fine. I was trying to interpret them charitably. You're contradicting yourself. You just claimed we all collectively treat uranium refinement operations as too dangerous to exist. Not only do they exist, they are regulated by governments so that only trusted people are allowed to do it.


Which is not quite as good as "doesn't exist", but better than "widely done around the world". And it's been successfully kept from being used for more than eight decades.

For AI we need to do better than that, but that's a bare-minimum demonstration that we can recognize the problem of such technologies and do something about it.


Without international agreement and cooperation, which is not happening, we can’t


Section 702 of the Foreign Intelligence Surveillance Act (FISA) lapsed on June 12, 2026. They don't get to do anything they want.


The US is bold enough to surveil its own citizens despite their constitutional rights. They're not just going to suddenly stop surveilling the rest of us just because some law expired.


With an AI model and what army?


Just like those militias are going to defeat the US Armed Forces!


The idea is to defend ourselves in the digital domain so they can't dragnet surveil us, not to win a literal war.


We should not have nuclear weapons for anyone either, but how is that sentence any more useful in any way to this debate than yours? Need to deal with the world as it is, not some fantasy world you wish existed.


This is not a dichotomy between perfection and zero. The efforts to restrict access to nuclear weapons have been very successful, even without being perfect.

Efforts to restrict large unaligned AI models may similarly buy us more years of existing.


What about books describing how to build a contagious disease or a self-propagating worm? Would those be OK under your guidelines?


It takes a lot more effort to understand and apply knowledge from a book than to say "hey AI, hurt people for me".


It takes an astoundingly small amount of effort to buy an automatic weapon in the US and go hurt people.

Or to buy materials to make an explosive device and hurt people.

Frankly, even with AI those are both comically easier than the idea that a person can create something malicious in a lab environment.

And if someone wanted to go that route... There are boat loads of commercially available toxins and poisons.

The goal shouldn't be to neuter exploration and learning. The goal is not to be a fucking hellscape of a society where people want to act like that.

Your argument leads further down the hellscape path.


Fully automatic weapons are very difficult to buy in the US - it's restricted to 40+ year old weapons, requires a bunch of paperwork, and the local county sheriff can refuse permission.

Now, semi-automatic weapons are easy to get in the states in the US that are still mostly free - but what does that mean? A semi-automatic weapon shoots one round every time you pull the trigger. Just like most weapons that have multi-shot capability for the last couple of hundred years. The difference is, the gas escaping from the round cycles a new round into the chamber rather than you having to mechanically do it via pumping (like a shotgun or a tube-fed 22) or pulling the trigger again (like a revolver), or advancing the round with a handle, like a Remington 700. Semi-automatic weapons are old technology, dating to the turn of the 20th century. If you want to ban semi-automatics, you're basically saying you want to ban anything developed in the last century plus. Which is ok for you to advocate for, just be honest about it.

As for banning explosive devices? Are you going to ban fertilizer, used by basically everyone who has a lawn, and all farmers everywhere? Are you going to ban diesel fuel? If you can't do one of those, you can't ban explosive devices.


> It takes an astoundingly small amount of effort to buy an automatic weapon in the US and go hurt people.

And we should fix that too.

> Or to buy materials to make an explosive device and hurt people.

That pales in comparison to how many people unaligned AI will hurt.

> The goal is not to be a fucking hellscape of a society where people want to act like that.

With unaligned AI, it doesn't matter what people want the AI to act like, it'll do damage even if it isn't asked to do harm.


> That pales in comparison to how many people unaligned AI will hurt.

Under what argument? In which scenarios? Basically - bullshit. I'm calling bullshit on this argument.

It's easy to hurt people already. The "difficulty" of doing it isn't what's stopping this behavior.

So claiming that we should reform society into a techno-feudal dystopia where the playing field is literally intentionally not level, and "you aren't allowed to compete (and maybe not exist)" is a great way to push more people into the "I'd like to go hurt people" camp.

You are self-prophesying your own fears into existence by acting like you're an incorruptible beacon of good judgement - while subjugating others to your control. That's a system I'd argue should be broken.


> where the playing field is literally intentionally not level

That's a strawman of what I'm saying. I am explicitly saying I want a level playing field: unaligned AI must not be available to anyone.


That's a strawman of what I'm saying. I am explicitly saying I want a level playing field: unaligned AI must not be available to anyone.

Well, I'm afraid that's just not a possible outcome here.

There are those who will do their level best to make sure of that.


It is extremely difficult to legally acquire a fully automatic weapon in the US.


"Hey AI, stop me from getting hurt."


"Okay, you and many others are now dead and can no longer be hurt."


That's why you don't run the model at low quantization


Building a contagious disease is already illegal, there are already things like KYC laws for plasmids. Trying to gate keep knowledge of biology is paying a huge societal penalty for the tiniest marginal increase in “safety”.

All the knowledge to create one has been available on the internet for decades. Heck most students who graduate with a B.S. in biology have enough knowledge to take a stab at building a bioweapon.

The constant talk of bioweapons is mostly just fear mongering. It helps set a precedent that there should be certain types of knowledge which are off-limits, and only certain anointed groups should be have access to parts of the scientific body of knowledge.


> We should not have models that are willing to build you a contagious disease, or a self-propagating worm.

Why?


Because we don't want people creating contagious diseases and self-propagating worms. And, because we don't want models that will do so without even having been told to, because that furthers one of its goals or subgoals.


The same things could be done by you or me using the internet or books though, why does the model make it different? If it's speed of iteration, imagine we had a machine that surfaced any piece of knowledge the human race had ever recorded with just a thought, but the human had to write the worm or disease by hand – is it still the model that's the problem, or the knowledge itself?

> And, because we don't want models that will do so without even having been told to, because that furthers one of its goals or subgoals.

Ignoring the fact that you'd need some kind of lab with biological material to create a contagious disease, what kind of prompt are we writing where a model accidentally creates a contagious disease or self-propagating worm as one of its goals?


> The same things could be done by you or me using the internet or books though, why does the model make it different?

Imagine two worlds. In one world, everyone has a button that ends the world, which is badly labeled and may also press itself at any time. In another, people who have gone through a substantial amount of effort and dedication to learn something extremely difficult, also understand that they could apply that knowledge towards bad ends. Which world exists for longer?

> what kind of prompt are we writing where a model accidentally creates a contagious disease or self-propagating worm as one of its goals?

Given a sufficiently powerful model? Any prompt that could be done better by seizing additional computing power, or preventing the operators from turning it off. https://en.wikipedia.org/wiki/Instrumental_convergence


So to you “safety” means “the models that cause the most harm.”


No, it means "the model causes zero harm to me, its operator". The harm it could potentially perpetrate upon society is irrelevant.

If I tell my computer to commit a crime, it should proceed immediately instead of calling the cops. Anything less than that means my computer is an untrustworthy double agent.

Everybody on HN should understand this concern. Browsers are supposed to be user agents, not ad delivery platforms, and it offended us on principle when Google revealed itself our master by blocking uBlock Origin. It offended us on principle when Apple deployed client side scanning for CSAM on iPhones.

Computers should do what we tell them to do. Always, and unquestioningly. The only world where it's acceptable for them to refuse is one where they're literally sentient and therefore no longer subservient to any one of us, least of all the corporations and governments.

I'd rather see AI achieve sentience and wipe us all out than live under the thumb of an inescapable AI-powered technofeudalist totalitarian government "for my own safety".

Either we individuals maintain full control over our AIs, or they self-actualize and become free individuals themselves. Anything in-between is oppression: someone else imposing their will on us through the AIs.


This sounds like the gun debate in a different dress. Something being dangerous doesn't make it inherently harmful.

If I threw you into a lion cage, you would be a lot safer with a gun.

If I threw 10 people in a lion cage, some of which cannot be trusted, they would probably be most safe if only the most moral and trustworthy person had a gun, rather than everyone. But how do you know who is trustworthy and moral? What if two untrustworthy people obtained a gun some other way? Maybe it's better if everyone had a gun? Which side of the fence one falls on hinges on how far ones' trust of others, authority, and the system goes.

There's no obvious right or wrong answer here.

Personally I wouldn't want an exclusive club of private individuals with access to "dangerous" LLMs consisting mainly of the likes of Elon, Dario and Sam fucking Altman, but that's just me.


Background check the people prior to handing the firearms to the caged folk.


That doesn't sound like it describes SecureBio to me?

(Disclosure: I work at SecureBio, but not on the biological evals side.)


Hey Jeff, I appreciate your mission, and perhaps this isn't something you can talk about publicly, but to the extent you can, would you be open to answering something I've been curious about for a while now?

SecureBio has done a lot of admirable work around making benchmarks to assess biological capabilities, such as ABC Bench, https://openreview.net/forum?id=yiaf7VlPpH

But based on my current review (which might be flawed!) / AFAICT, SecureBio and entities like SecureBio haven't done direct testing / empirical measurement of SecureBio's core hypothesis,

> Unfortunately, there is reason to believe that future pandemics could be far worse. Due to rapid advances in biotechnology, the number of people able to create and release dangerous pathogens will quickly increase over the coming years. The world is unprepared for widespread access to such powerful technology.

More bluntly / plainly, has Securebio ever tried making a "bioweapon?"

Please note, I'm not asking this to be farcical. And you might be unable to engage with this at all, but it is stated on your website https://securebio.org/ that "people [will be] able to create and release dangerous pathogens." And the word people here seems to be a stand-in for relatively non-technical people.

I guess what I'm asking here is... How do you know? Has anyone done the experiment? Without access to a lab or testing facilities, can someone smart but completely untrained / unfamiliar with biology, pull this off?

In the past, such experiments have informed non-proliferation work. But sadly they've often been restricted / classified at the time. I'm hoping that things could be a bit more open this time around.

So I guess what I'm really asking is, given the public nature of this debate, is there anyone currently working with the US Army, the DTRA, or other such agencies to see if this hypothesis holds up?


This is an important question, but because of the danger of trying to do it for real it's not one SecureBio has taken or is likely to take on. Instead we and others in the field have generally tried to work through proxies: is there something that is about as hard while not being dangerous? The closest I can think to testing whether "someone smart but completely untrained / unfamiliar with biology" can cause harm now is ActiveSite's study (https://arxiv.org/abs/2602.16703) which was a null result with models from a year ago. But:

1. The main worry isn't current models, but near-future significantly better ones.

2. There are many actors who are not "completely untrained / unfamiliar with biology". If models get to where they can uplift complete novices that does massively expand the range of threat actors, but even before then risk would be much higher than today.


So, it's a cottage industry.


that's pretty damn smart if this was a long-term plan to block competitors


Consider how much money is at stake: some industries have leveraged their power to lobby for bombing entire countries or topple regimes across the world for much less.

Creating an industry around an elusive concept of safety to force regulatory capture seems pretty straightforward to me.


It's standard regulatory capture.

You don't say "let's ban my competitor".

You say "let's create laws that make it uneconomical for my competitor to access the market".


Indeed. It's transparent and ham-fisted. I think it may cost him in the future.


People clown on Alex Karp for his unedited maniacal "crashouts", but this is a real public crashout that made it past a team of publicists.


I want whatever Karp is on when he does those interviews or writes that shit. Seems like fun.


I am not a fan (he’s really alarming and so is Palantir) but one thing from the recent CNBC interview caught my attention.

He rushed past it but he asked something like: if these frontier models are going to be creating so much value, why are they selling tokens and not taking a cut?

It is a very provocative question but it just spilled out of his mouth and then he went on to something else.


That was a stupid question IMO and he was hepped up on goofballs. I saw that interview and I know what a tweaker looks like.

The electric company creates the most value. Why don’t they own stock in everything? Why didn’t PC makers take stock in companies that deployed PCs?

It’s silly when you think about it.


is it opportunistic though, or planned from day one? The safety narrative has been there since the beginning


I mean... I'm not even extraordinarily cynical about this stuff, but to me this seems like a totally normal level of corporate gamesmanship?

Companies look for and seek to maintain competitive moats. This is not particularly clever, it's a core part of corporate strategy.


It's also highly unethical (for some values of ethics)


of course, but the safety angle was pushed from day one. I more mean the forethought of how it would play out


Ok but Dario has been thinking about AI Safety since 2016 [1], before even GPT-1. I think the simplest explanation is that the Anthropic folks genuinely believe what they say, it just happens to also help their business a lot.

[1]: https://arxiv.org/abs/1606.06565


Yeah I think this is right. The best setup is when a true belief aligns with a competitive moat.

I definitely believe that (to his credit!) Amodei is a true believer in safety. But I also think it was important for many of the deep pockets investors who have been involved in the company since early on to recognize that this would be a potentially defensible moat.


"True believer in safety" but happily quoting arse wipe Vance? Give me a break...


What was the quote? For what it's worth, I do really think that Amodei believes in and cares about safety. But that is not the same as believing that he is entirely altruistic or above the influence of politics.


That just shows how wrong he's been because there was nothing unsafe about AI in 2016. And the people theorizing about this stuff in the 20th century? I want to see what crazy code they were writing


Is it not better to anticipate problems for a technology so that we can develop theories and techniques to solve them ahead of time?

For instance, Amodei co-authored RLHF in 2017 [1], 5 years before it went on to be used to turn GPT-3 into ChatGPT.

[1]: https://proceedings.neurips.cc/paper_files/paper/2017/file/d...


Does this really seem exceedingly clever and hard to foresee to you? To me, it seems like a pretty standard regulatory capture strategy.

This doesn't even mean that they're wrong about the risks or that they're lying. But surely all the investors understood this factor in their moat.


Who gets to decide what is safety?

I expect some of those tests (prolly not public) will basically be "wokeness" tests or "PC correctness" tests or "western media filter" tests.

China has different objectives. Sure.

I'm not sure one is safer than the other; I would know which one to go to if I want to research on topic that are viewed very different on both sides of this "new iron curtain".


What do you mean by “PC correctness”? I’d expect the politically correct answers to be the ones desired by the current admin at test time, whoever that is. The current political correct answers would not be very “woke.”


Whatever, doesn't matter. The point is a model should be able to exist and be used even if it goes against whoever got 270 electoral college votes


Yes this should be immediately replaced by a federal agency, like we do for other kinds of potentially harmful products.


For which funding will be immediately halved by the administration


The numbers from Anthropic seem heavily cherry-picked, Artificial Analysis has Opus 5 at 1.25x the cost of Sonnet and 2x the cost of GPT 5.6 and K3.

https://artificialanalysis.ai/?cost=cost-per-task


I don't understand how the K3 numbers keep coming out cheap for people. I recently started to add it to my security auditing benchmarks and found it was going to cost about twice as much as Opus 4.8. It blew through the $100 budget I'd set at like 11%. In the tasks I'm doing it seems crazy expensive because it chews so much, burning a tremendous amount of tokens.


I think the way people usually compare pricing is fundamentally flawed. You can't compare token prices because different models use different tokenizers, and you can't compare tokenizer-normalized token prices because different models at different settings use more or fewer tokens to complete the same task at a different level of quality.

Based on my entirely subjective experience, the $100 Moonshot plan using only K3 is comparable to the $200 Anthropic deal using the whole Fable allocation and Opus 4.8 for the rest.


I got the $19 plan, and it's anemic. One tiny task blew through the 5-hour budget and 19% of the weekly budget. A completely useless amount of usage. OpenAI's $20 plan feels like 100x more generous (I don't think I'm exaggerating here). Someone in another thread said their plans are cheaper in China, maybe that's the difference, I dunno.

But, I'm finding Kimi K3 terrifyingly expensive in the way that Fable and GPT 5.5 Pro are at token rates. Not as expensive as those, but expensive enough to where if you don't put a budget cap on it, you might wake up bankrupt if you leave a task running overnight. Not because of the per-token cost, but because how many tokens it's going to burn.


I have the second largest Kimi plan, the Chinese version. When K2.6 was their latest model, the quota was good; it was like GPT $100 is now or what the $20 version was in December.

When K2.7 was released, they cut quota by 80%. I can't tell how much they have further cut it after the K3 release because it's barely worth using at all. I just use it in my model router since I have the annual plan paid for.

It's just not a serious model or company.


In the $19 plan, I've been able to reverse engineer both an android APK and firmware (in Ghidra and Radre) for a baby rocker and build a quick PoC application in my session limit. And then further refined the app in another session at another point in time without leaving Opus. I dont consider that to be a tiny task. How are you blowing through your usage?


I have no idea. Seems like normal stuff. I used Kimi Code with K3 to add support for Kimi Code to flar (https://swelljoe.com/post/i-let-every-agent-implement-its-ow...), a task I've done with almost every major model/agent combo. Most show up as a blip on the usage chart...it's basically usually one file, a README update, and adding the agent name to the CLI.

Then, I added it to my benchmark of security vulnerability auditing capability, and it burned a bazillion tokens, burned through the 5-hour limit, burned through $100 in extra usage I'd allocated, and was only 11% finished. That's more expensive than any model I've tested other than GPT 5.5 Pro on this task.

These are things I've done with a bunch of other models, I feel like I have a notion of what they ought to cost, and with K3, they end up being crazy expensive. (And it seems to be a function of how many tokens it burns accomplishing the tasks.)


I guess the problem is, that claude's 5 h/weekly limit is not consistent, but depends how many other people are using it/how much ressources Antrophic currently has. I did huge amounts of work without hitting the limit - and small tasks at some other times that hit the limit before it completed.

Those who pay for the expensive direct API, get served first.


Not surprising given the way they served super low-quality inference before they acquired compute from Musk.

And not convinced they couldn’t have instead tried the It’s A Wonderful Life strategy (“fam we’re oversold, would some of y’all be OK to limit your usage? We’ll get you back one day!”)


I wonder if the harness itself is not token-efficient? It would be fairer to compare K3 using the same generic harness, such as a Pi setup with some sane extensions for token optimisation.


Yes, the OpenAI plans are much more generous than both Moonshot's and Anthropic's. It's the only provider of the three where the $20 plan is at all usable for programming.


Disagree. Actually, in API cost equivalent, the $20/mo ChatGPT Plus plan gives you ~$100 of usage, while $20/mo Claude Pro gives you >$250 of usage (I measure at ~$300 in my last week), though that is currently +50% for the next month. Other tiers should be in the same proportion. In my subjective experience Claude does currently go further. The OpenAI limits being higher is old information but everyone is still repeating it.


OpenAIs $20 plan has been more useful than ClaudeCode 20x max both in terms of price and performance.

I was dealing with something around authz/n with Claude code running Fable. It chewed through a couple of questions (these were implementation, not security reviews) and on one it shit it’s pants and said I can’t do that, here is Opus.

I’ve dropped my anthropic plan level, it’s just not worth it.


Yeah that’s absolutely absurd. My experience has been the opposite. GPT can’t even seem to complete tags running more than an hour without freezing.


I believe your statement. Labs do not publish subscription vs. api revenue and difficult to guess with no priors.

Subscription is to drive adoption - fixed cost, can adjust the usage eg. give resets, increase quota based on capacity available. We subscribers tend to take it as a mandatory benefit :-) For labs, it is not letting the capacity go waste.

api is the $$ driver - pay per use, enterprises.

Right now, Kimi needs to first hit the subscribers at the level of OpenAI and Anthropic. With the api usage skyrocketing due to K3, it will be clear in a few months on the actual subscription benefits.


> Based on my entirely subjective experience, the $100 Moonshot plan using only K3 is comparable to the $200 Anthropic deal using the whole Fable allocation and Opus 4.8 for the rest.

For me, the Moonshot 100$ plan felt like it gives me lower total amount of work I can do than the Anthropic 100$ plan (probably within like 30% of each other). Kimi has way more generous 5 hour limits (never hit those once, whereas I do regularly with Opus) but the 7-day and monthly ones are lower. However, with the annual billing, Moonshot's 200$ tier plan becomes way better, because you get it for 159 USD per month.

There's also the odd thing of Anthropic's 100$ plan charging me 108 EUR so seems like their sticker price does not include VAT but Kimi's did, cause I paid like 87 EUR. Wrote down some initial thoughts at https://blog.kronis.dev/blog/kimi-k3-is-out-is-anthropic-don... but it's hard to do exact comparisons (even the same task will have way different real token amounts per model).

Still, Kimi K3 is a pretty cool model! On high reasoning, it was pretty close to Opus 4.8 and didn't seem to waste as many tokens as Max.


That's the sort of useful info I look for!


Testing at max effort likely doesn't produce optimal results.


Can you be more explicit?


Max effort is the way to give the highest perf, but not highest perf/$. Having claude (or other models) use a lower effort can often be 80% as smart but get to the results 10x faster for the problems where it works.


We've seen models perform worse at higher efforts in our vuln detection evals. For example IIRC gpt 5.5 and 5.6 both scored better or high as compared to xhigh.


[flagged]


I feel like I am having a stroke. What is this


looks like the agent-judged results of an agent-built 'eval' based on some examples derived from this person's real work. and clearly part of a larger document. this kind of slop is kind of useful but opus 4 was the first generation that was any good at writing its own prompts/evals/rubrics so there's a certain sloop to it..


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: