Hacker Newsnew | past | comments | ask | show | jobs | submit | nr378's commentslogin

The competition for the least reliable developer service continues between GitHub.com and Claude.com...

GitHub needs to completely bifurcate their enterprise/paid services from their free services at the infra level.


They have that-ish as an option: https://docs.github.com/en/enterprise-cloud@latest/admin/dat...

I'm told that GitHub has asserted to us that moving to this model means we would not be exposed to github.com outages. It's not at feature parity with github.com though.


Thanks, this option is good to know.

We're currently on "GitHub Enterprise Cloud" on github.com and are affected by this outage (even though we use self-hosted runners!), but we're not on "GitHub Enterprise Cloud with data residency" on *.ghe.com, which I understand is/may not be affected by this outage?


This is what they've told us. It's represented as basically a separate deployment of the entire GHEC stack, so you're not exposed to the load/scaling issues they believe are the underlying cause of all the github.com outages today.


The last large company I worked for switched from self hosted github enterprise to github.com in January 2025 or so.

I wonder how much egg is on that exec's face.


Do you have any meaningful level of faith in GitHub's ability to deliver on stability? At this point, I have none.


Meaningful is subjective, but yes I do. It was very stable for many years, and I do believe the recent issues are mostly or all because they were caught flat-footed by the rapid AI-driven load increases.

This has been a bad time, though. I'm ready to move back to self-hosting if they can't get it together or we move to GHDR and it's still bad.


A lot of the stability problems come from trying to scale a free product on a still WIP cloud solution without costing too much; the same software running on separate, paid for infra has a lot better odds.


According to their status pages (e.g. https://eu.githubstatus.com/, https://us.githubstatus.com/), their Enterprise Cloud uptime for Actions is significantly higher.


“GitHub Enterprise Cloud with data residency” is hosted on separate infrastructure and dedicated subdomains under *.ghe.com. It’s been around since November 2024z

It’s not the same thing as GitHub Enterprise Cloud hosted on the shared global network on github.com.

https://docs.github.com/en/enterprise-cloud@latest/admin/dat...


So confusing and so Microsoft. They love to have licensing so complicated their own sales people aren't up to date and have to rely on third party spreadsheets.

Edit:

Read that document - why do more people not do the self hosted option with GitHub Enterprise Server ?


At the point you are self hosting, you have many more options ranging from a simple ssh git server with pick your favorite cicd, to gitlab ce, to forgejo, and more.


That is a different and later product with a confusingly similar name.


Just to be clear, I am on Github Enterprise, and am also experiencing this disruption both privately and publicly on every org and project I have access to.


That's what I don't understand. They could mitigate their name so much if they just split free/paid/enterprise. It's already shown that enterprise is much more estable and is largely unaffected from service disruptions. Why don't they go one more layer? For sure it's worth the extra complexity.


There is no such thing as "just split" there is 20+ years of legacy decisions and even if the split is relatively clean it is still probably 1 years work for 200 people for maybe a marginal improvement.

The real money is going to go towards, "make this all more reliable".


Depending on the cause of the current issues, that move would likely cause more harm to paid services than good.

Their last postmortem made clear that their challenges are operational. Scale puts pressure on operation, but it's not what blocks them from keeping up.

Doubling the operation doubles the operational challenges.


That's what Azure DevOps is supposed to do, but for some reason GitHub has a redundant enterprise division.


surely if they did that everybody would complain how github "lost its touch with open source since they now prioritize paid services"


enterprise is mostly separate, is it not? uptimes are significantly more reasonable on the enterprise status pages


We are in GHEC right now and GitHub Actions is not working. It's been down every time githubstatus.com says it's down.


Same for us, I'm not even sure what product that other "Enterprise" status page refers to..


What country did you choose to host your data in ? Could be region based


It would probably be better to run projects with extremely high commit/merge frequency on a separate "slop infrastructure", basically like MMOs move cheaters to their own servers ;)


High growth companies often have significant negative cashflow during the early high growth era, followed by positive cashflow in the years later down the line.

This phenomenon is known as the J-curve[1], and Uber is a good example of how this can turn out absolutely fine. To some extent, the entire Venture Capital industry exists to finance precisely this dynamic!

Nb. I'm not suggesting OpenAI is fairly valued, or that they will definitely become profitable, but "OpenAI is losing billions of dollars" doesn't really mean anything in and of itself.

[1] https://www.uark.vc/blog/breaking-down-the-j-curve-the-journ... (many other similar such articles exist)


> High growth companies often have significant negative cashflow during the early high growth era, followed by positive cashflow in the years later down the line.

Uber is the antithesis of OpenAI, it’s not a good example. Uber was burning money on acquiring customers. OpenAI is burning money to provide their service (and the R&D they need to continue to have valuable models). They cannot just stop and turn profitable like Uber. The money they burn isn’t invested, it won’t yield a multiple of revenue in the future. It’s consumed for compute and that’s it loo


If the leaked data is to be believed, OpenAI is spending 40% of revenue on sales and marketing, which is not the OPEX profile of a product-led technology company

Broadly speaking, companies that spend 40%+ of revenue on sales and marketing end up being a bit of a drag on society. Eg. Salesforce’ product quality is far lower than winners in other sectors that sit closer to 10-15% of revenue on sales and marketing

Maybe - just as how the city of Sao Paolo implemented a ban on billboards - we can implement a law where a 3 year rolling average of sales and marketing spend cannot exceed 20% of revenue in that period


> They cannot just stop and turn profitable like Uber.

Of course they can. They could just stop training new models and milk the existing ones. A billion users check in ChatGPT weekly. Software developers wouldn't stop using Codex.

OpenAI is not unlike any other startups who try to build their marketshare early on. No matter how much money they lose, they would be fine as long as they could raise more money than they spend. Uber is exactly the same. HN during 2015-2020 were full of comments predicting Uber's demise.


I don’t think you understand how bad OpenAI economics are. The company is burning billions just to operate. They cannot stop the training treadmill due to competitive pressure, but assuming they do that would only reduce their expanses, not increase their revenue. They would still be in the negative. We are talking about a company that has more than >$750B of infrastructure expenditure commitment for 2030.

The number of users they have checking weekly is irrelevant, most of them are free users, unless they find a way to make money from them, but their ads business has been a flop so far.

For context: Uber losses were $12B over 5 years. AWS was $5B invested over 7y.

OpenAI is projected to lose more than $14B just this year!!!


> Uber losses were $12B over 5 years.

Uber burned through roughly $32 billion in cumulative losses before reaching sustained profitability.

The rough timeline:

- Founded 2009, and lost money every year for about 14 years

- Biggest single-year losses: ~$8.5 billion in 2019 (the IPO year) and ~$9.1 billion in 2022

- 2023 was its first full year of net profitability, earning about $1.9 billion

- Uber has a market cap of $153bn as of today (at a P/E of 16.5)

OpenAI has received substantially more funding than Uber, so its losses will be substantially higher (spending investor money shows up as a loss on your P&L), but again that doesn't mean anything in and of itself.


> The frontier LLM labs run on a huge fixed cost and very low marginal cost.

> Imagine that you want to buy a few B300s to run GLM 5.2 and rent the service out to other people. How could this business be viable and sustainable in the first place?

My understanding is the frontier labs have huge fixed costs and relatively low marginal costs because they have to bear the cost of training the model/R&D, and then amortise that cost over their userbase.

By contrast, if I buy a few B300s and run GLM5.2 and rent the service out to other people, I can be profitable at a comparatively very small scale because I got the model for free.


Arguably Google is both (with GCP and Gemini).


From GP's point of view, Google is not a data center company because they're not renting out data centers to other companies; they're using them for their own usage.


Claude Teams and Claude Enterprise are 2 distinct plans. Simon is right that Enterprise seats have no included usage (and so all usage is charged at API billing rates), whereas Teams seats do.


The data doesn't well support the claim that FP is best. Elixir tops the table at 97.5%, but C# (88.4%) is OOP and scores almost identically to Racket (88.9%), and Ruby (81.0%) and Java (80.9%) both outscore Scala (78.4%), which is explicitly functional. If FP were the driver, Scala should beat those languages, but it doesn't.

It's tempting to argue that a more constrained language helps, but Rust (62.8%) vs Elixir (97.5%) is an interesting data point here. Both are highly constrained, but in different directions. Elixir's constraints narrow the solution space because you can't mutate, you can't use loops, and you must pattern match, so every constraint eliminates options and funnels you toward fewer valid solutions that the LLM has to search through. Rust adds another constraint that must independently be satisfied on top of solving the actual problem, where the borrow checker doesn't eliminate approaches but adds a second axis of correctness the LLM has to get right simultaneously.

Overall, it seems like languages with strong conventions and ecosystems that narrow the solution space beat languages where there's a thousand ways to do something. Elixir has one build tool, one formatter, one way to do things. C#, Kotlin, and Java have strong ceremony and convention that effectively narrow how you write a program. Meanwhile JS, Python, PHP, and Perl offer endless choices, fragmented ecosystems, and rapidly shifting idioms, and they cluster at the bottom of the table.


Scala is explicitly multiparadigm and offers a lot of advanced OOP features. It also had a Python-like (though reportedly better handled) 2 -> 3 transition, which deprecated some things, removed others, and added a bunch of new ones. Scala has always been complex, and right now it's also chaotic. It's a wonder the models can get that high a score with it, honestly.

Racket is a similarly large PL, with many abstractions built on the metaprogramming primitives it offers. Without looking at the generated code, it's hard to say anything, but I suspect the high score despite that might be because of the Scheme core of Racket: `racket/base` is a much smaller language than `racket`, so if the LLMs keep to it, it might narrow the solution space enough to show different results.

In general, I think you're half-right: the "solution space" size is a factor, but so is its shape - ie. which features specifically are offered and how they interact. A more compact and cohesive language design should yield better results than just a reduced surface area. C is not a huge language, but the features it offers don't lend themselves to writing correct code much. Elixir is both relatively small and strongly steers a programmer towards safer idioms. Racket is big, but the advanced features are opt-in, while the baseline (immutable bindings, pure functions, expressive contracts) is similar to Elixir. Python is both huge and complex; "there's one obvious way to do it" has always been a bit of a joke. Rust is incredibly complex - the idea is that the tooling should allow you to handle that complexity easily, but that requires agents; one-shotting solutions there won't work as well.


What if it is the quality of data? Internet is full of terrible python/js, but probably not Elixir.


Seems plausible. I used to refer to StackOverflow before LLMs and a good amount of the examples there were flawed code presented as working. If the LLM had less junk in its training then it might benefit even though the volume of training on that language is lower.


If we assume that the amount of training data matters at least a bit (which is a very reasonable assumption), I wouldn’t immediately discard the functional hypothesis. Scala’s score is almost equal to Java’s even though there’s probably something like two orders of magnitude less Scala than Java code in the wild. Similarly with C# and Racket.


Yep I think you can reasonably argue that immutability + strong conventions are the most important dimensions (as opposed to FP vs. OOP, as much as I like FP and dislike OOP):

Immutable by convention + Strong conventions: 91.3% - Elixir 97.5%, Kotlin 90.5%, Racket 88.9%, C# 88.4%

Immutable by convention + Fragmented: 78.4% - Scala 78.4% (n=1)

Mutable + Strong conventions: 77.5% - Ruby 81.0%, Swift 78.5%, Julia 78.5%, Dart 78.0%, Go 71.7%

Mutable + Fragmented: 67.9% - Java 80.9%, R 75.8%, C++ 75.8%, Shell 72.9%, Python 65.3%, Perl 64.5%, TS 61.3%, JS 60.9%, PHP 53.8%

(my grouping is somewhat subjective)


I agree with you, but, from the article: "The amount of training data doesn’t matter as much as we thought. Functional paradigms transfer well"

Anyway, I tend to think you are right, and the article is wrong in that sentence. (Or I misinterpreted something?)

I think both the quantity and quality of that has a big influence in the results.


I took that to mean ≈ "Amount of training data isn't the big factor dwarfing all else." Depends who "we" refers to, I guess. Back when LLM-generated code was new, I definitely saw predictions that LLMs would struggle with niche or rarely used languages. These days, consensus among colleagues within earshot is that LLMs handle Rust much better than Python or C++ (corpus size and AutoCodeBench scores notwithstanding).


TFA's theory also doesn't explain why C++ (75.8%) beats Python, JavaScript and Rust.


> 3. Storing it the way this article presents makes it usable for agents, but not humans. Whereas the point of knowledge graph, ontology, etc is to create the same layer for both humans and AI to interact with

If storing it this way makes it usable for agents, then why don't humans just use agents when they need to interact with it?


Let's say that you want to know who your largest customer is, both by order value and volume. I could either: 1. Prompt my agent and deal with writing the prompt, waiting for the agent to sift through all the data (which would be massive), and pay the token costs, all of which has to be repeated everytime I want to answer this question, OR

2. I check my ontology for the answer, probably in a dashboard, and it takes 5 seconds. I have a link I can freely share around my enterprise and I haven't spent token costs.

Whats more, when I have sent my agent out to some tasks (go find out what revenue we're leaving on the table by not selling spot contracts to our biggest customers) my ontology gives me a few bits of data to validate the agents work against. For humans and AI to work together, they need the same context layer


Dario has made a specific cohort argument here. His numbers (from various interviews) are: you train a model in 2023 for $100M, deploy it, and it earns $200M over its lifetime. Meanwhile you train the 2024 model for $1B, which goes on to earn $2B. Each vintage returns 2x on its training cost.

However, the GAAP P&L tells the opposite story. You book $200M revenue in the same year you spend $1B training the next model, so you report an $800M loss. Next year you book $2B against $10B in training spend, reporting an $8B loss. The business looks like it's dying when every individual model generation actually generates a healthy profit.

That's actually Dario's answer to your depreciation question. If each cohort earns back its training cost within its natural lifespan (however short that lifespan is), the depreciation schedule is already baked in. The model doesn't need to live forever, it just needs to return more than it cost before the next one replaces it. Whether that's actually happening at Anthropic is a different question, and one we can't answer without audited financials, but it's the claim Dario makes (and seems entirely reasonable from a distance).


GAAP doesn't work here really. the R&D treadmill means you are always betting on next year and its NOT inventory or something you can defer your cost on. It's an upfront R&D expense.

so what happens on year 10 when Anthropic hits a $10B training and only returns $8T? they're cooked


Yeah, that's kind of what I'm wondering about.

It's an interesting story about how even though all metrics show massive losses actually they have massive gains.

Accounting is a rather mature field, so I figure that someone in the past has tried this stunt and there should probably be ways for dealing with it.

Or do they always flame out after losing all the money? Knowing the history here would be informative.


If those numbers are correct, then my assertion that "Almost certainly, any reasonable depreciation schedule of the cost of training will result in leading labs being presently wildly unprofitable." is incorrect.

And I admit that I made that assertion from my gut without actually knowing if it's true or not.


If you have to continually spend greater amounts of money to keep up with the competition on every new model then it is dying.

Every single time a company comes around and goes "Actually GAAP are wrong, look at my new math that says were good" its led to much wailing and gnashing of teeth in the future when it inevitably isnt.


That's an interesting idea. I'm curious, though, are there any other industries and/or companies that have tried to pull this sort of thing off? And what ultimately happened to them?


Enron had a system like this. They regularly worked on large, long term contracts that became profitable over years/decades. They wanted to push rewards forward so would estimate the total value of the contract and book the profit when it closed. Mark-to-market accounting wasn't unheard of the time but using it for assets without an active market was unique. Without the market to make against, the numbers were best guess projections.

The problem is everyone along the line is incentivized to be aggressive with estimate (commissions for sales are bigger, public financials looks better) and discouraged from correcting the estimates when they go wrong.

Estimating multi-year returns on frontier models looks harder than estimating returns on oil and gas projects in the 90s.


The bar for "wildly unprofitable" has risen quite a bit since then, but Amazon basically pioneered this.


Why would anyone use 200M model when 1B model is available? The company increase its bet with each iteration increasing risks. It blow up at some point because it cannot guarantee 2B return after 1B investment.

To GAAP point - 200M or 1B or 10B is not a loss but cash converted into an asset. It won’t affect the bottom line at all. Unless the company re-evaluates the asset and say it now cost 1M instead of 200M. This would hit the bottom line.


If you can remember where you read it, could you share a link?


https://youtu.be/GcqQ1ebBqkc?t=1027 is on such but he doesn't actually say that each model has been profitable.

He says "You paid $100 million and then it made $200 million of revenue. There's some cost to inference with the model, but let's just assume in this cartoonish cartoon example that even if you add those two up, you're kind of in a good state. So, if every model was a company, the model is actually, in this example is actually profitable. What's going on is that at the same time"

importantly you'll notice that he's talking revenue, and assumes that inference is cheap enough/profitable enough that 100M + Inferance_Over_Lifetime < 200M


Based on the docs and API surface, I think the filesystem abstraction is probably copy-on-mount backed by object storage.

I suspect it works as follows: when a task starts, filesystem contents sync down from S3/R2/GCS to a local directory, which gets bind-mounted into the container. The agent reads and writes normally - no FUSE, no network round-trips per file op. On task completion or explicit sync, changes flush back to object storage. The presigned URL support for upload/download is the giveaway that object storage is the source of truth.

This makes way more sense than FUSE for agent workloads. Agents do thousands of small reads (find, grep, git status) that would each be a network call with FUSE. With copy-on-mount it's all local disk speed after initial sync.

Cross-task sharing falls out naturally - two tasks mounting the same filesystem ID just means two containers syncing from the same S3 prefix. Probably last-write-wins rather than distributed locking, which is fine since agents rarely have concurrent writes to the same file.


That's a good analysis:) We want to go with FUSE but the performance overhead, especially with multiple calls to use files, is a constraint


How have you determined that? You can easily push 6GB/s+, sub ms ttfb with networked filesystems, and hundreds of thousands of iops through fuse.


sprites.dev / fly.io has publicly said they are using a variant of JuiceFS for the object-storage-to-VM-filesystem stuff, it's cool tech.

* https://fly.io/blog/design-and-implementation/ * https://juicefs.com


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: