"postgres is so stable I will never trust a rewrite."
"covering 100% of postgres regression suite doesn't guarantee you have replicated every behavior."
Folks who are okay with LLMs think the regression test suite is the spec and is the guarantee of stability. How else can it be? If you are depending on some behavior not covered in the regression tests, how do you know the next minor release won't break you?
Folks who are against seem to imagine a platonic ideal of PG which conveniently is the original PG implementation by tautological definition. So no rewrite can ever meet their bar.
> Folks who are okay with LLMs think the regression test suite is the spec and is the guarantee of stability. How else can it be? If you are depending on some behavior not covered in the regression tests, how do you know the next minor release won't break you?
I think this deserves a real response.
First, an analogy. I drive along a cliff with no guardrail. How do I not drive off the cliff? By knowing how to drive. Sometimes people mess up and drive off the cliff
In practice people are operating Postgres as a machine, more than an abstract spec. Minor releases exist, but the changes are made by people operating the machine and who have a fear of making the machine break.
There are also performance characteristics that are part of an informal spec. While you can definitely write regression tests on performance, theres loads of value in stability of internals because people who have problems look into the machine and discuss it.
The internals might change, but there’s a lot of friction. So… you can have a lot of informal knowledge about.
If everyone working on a software stack is bathing in this informal knowledge, then the decision making is based on that. Things like what is meant to be in a minor or major release is understood. And people make those judgement calls.
After all, even if you have regression tests if you’re making changes you’ll need to write new tests! How do you know your new tests are right? That the new behavior is right?
PG is the entire machine. Fortunately we have version numbers, migration strategies, etc. But “here’s a new binary that passes the test suite and… maybe changes everything not covered by tests maybe doesn’t”. Why bother suffering when there are totally reasonable gradual rewrite strategies?
And of course… every bug in the world… got through despite a test suite! What software out there doesn’t have bugs?
And if the behavioral difference in the rewrite does affect people… I guess that’s something right? “Oh this isn’t a bug because it wasn’t covered in regression tests” is not that tenable.
If you are saying the PG regression suite doesn't cover these perf tests - that is fair. I consider perf to be part of the "regression framework" generally.
>> How do you know your new tests are right? That the new behavior is right?
This is a more generic question in the LLM world. Reliable verifiers is what drives LLM loops. If you don't have these, you have no idea if you are actually making progress or you just generated code that returns 42 for every question. You needs something to actually ground the LLM output against. For existing features, the reliable verification is the existing regression/perf suites. For new features, the regression/perf suites should expand to fit. There are ways to go at this that range from ad-hoc (line/branch coverage) to fully formalized (verification-aware languages like Dafny).
> If you are depending on some behavior not covered in the regression tests, how do you know the next minor release won't break you?
Decades of intrinsic knowledge. Which a rewrite lacks.
imho the issue is the language used. AI rewrites are cheap and cheap exercises require a great deal of scrutiny and proof of correctness. Simply using regressions test lacks the intrinsic knowledge of a decades old codebase stuck in developers minds.
Claiming a rewrite is better because it passes all tests is a flex, a new version requires a boatload of evidence for people to accept it as an improvement and not just "passes tests, written in rust via LLM so it must be better". Run it in production for a year in a sufficiently large system and you might be somewhere.
>> Decades of intrinsic knowledge. Which a rewrite lacks.
You mean the ones encoded in the regression suite? I think your argument is valid in many medium-to-faang firms where the application is going to have encoded business logic that isn't explicitly tested for. If anything, projects like PG are the exact opposite: no business context and regression suites that test every possible scenario due to the accumulation of bug fixes and context over many years.
>> Run it in production for a year in a sufficiently large system and you might be somewhere.
How do you know every release of PG doesn't break in the setting of "a year in a sufficiently large system"?
> no business context and regression suites that test every possible scenario due to the accumulation of bug fixes and context over many years.
Spend time enough with a codebase and you’ll know stuff about its behavior that is not encoded in a test suite. Especially when you need to adjust an integration tests due to the modification of an invariant in a dependency.
> How do you know every release of PG doesn't break in the setting of "a year in a sufficiently large system"?
Because the postgres team is professional and will take care of publishing a changelog for what has been modified since the last version.
I think what folks want, but aren't quite able to articulate, is an ongoing community and effort that indicate a project will be healthy and maintained. We want to be able to rely upon the software that we are choosing to use.
Regardless of the technical choices, whether Rust is better or worse, whatever -- pgrust popped into existence thanks to one person driving an LLM through 7000 commits in ~2 weeks. It produced something that passes the regression tests. Even as an LLM-sceptic, I think that's amazing.
From that point, though, it appears to have been completely abandoned. There hasn't been a commit in a month, other than a brief tweak and a note that an as-yet unpublished version that's even betterer is in the works. IDK. I don't think we've acclimated to the shock of the change LLMs create, but if the outcome is a forest of exciting new projects that have a bus factor of 1 and little to no collaboration, I think that's a disservice to this profession.
I am not saying anything about the viability of that particular project; just that one argument that was repeatedly made in that thread which I think is frankly nuts.
Yeah. I guess what I'm saying is that I feel like we're all still stuck debating these things on technical merits alone. The `bun` rewrite and `pgrust` both expose something -- I find, at least! -- uncomfortable about how we understand the technical side of our profession, but I think the social side remains the same.
You cannot access Fable because Anthropic can't reliably tell whether you are a US citizen. The govt order is based on export controls to non-US citizens.
You can already imagine Anthropic working with a bunch of shady brokers to "remedy" this situation.
This particular order wasn't actually about citizenship at all. It seems the administration simply believed restricting the order to non-citizens would make it easier to defend in court, but they made it knowing full well that the only way to implement it would be to completely shut off access for everyone.
>> No matter how well “AI” works, it has some deeply fundamental problems, that won’t go away with technical progress.
Without an explanation of what they author is calling out as flaws, it is hard to take this article seriously.
I know engineers I respect a ton who have gotten a bunch of productivity upgrades using "AI". My own learning curve has been to see Claude say "okay, these integration tests aren't working. Let me write unit tests instead" and go on when it wasn't able to fix a jest issue.
A part of the job is only enabled when you get the Principal label. Unlike almost all other transitions, you only prove that you can do the role when given the opportunities. The hardest part about this transition is that you are doing two almost orthogonal roles - Sr. SDE / Tech Lead and the principal parts. It is very easy to not show impact in the former while chasing the latter.
A part of the job that is completely different to what you've been doing, only you were promoted to it because of what you've been doing?
Isn't this like a recipe for the Peter Principle?
"The Peter principle is a concept in management developed by Laurence J. Peter which observes that people in a hierarchy tend to rise to "a level of respective incompetence": employees are promoted based on their success in previous jobs until they reach a level at which they are no longer competent, as skills in one job do not necessarily translate to another"
Nobody performs as the CEO until they are given the CEO title. If the Peter principle was true 100% of the time, we wouldn't have any successful CEOs ever. Which is clearly not the case.
Some CEOs do, some don't. Some are successful, some are middling, some are shameful failures or frauds.
However, none of this has much to do with what I said: isn't promoting someone for things they would have to stop doing (as in TFA) a recipe for the Peter principle?
At some point, we have to understand articles and blog posts like the one we're discussing are mostly fluff; an ad for the person writing them. They are self-promotion, there's not much sense in trying to extract valuable lessons from them.
----
I thought of another analogy:
"You're an excellent marksman and a sniper, therefore we promote you to... general of the army!"
But, you could argue, maybe the sniper was already directing strategy, and that's why he got promoted? Nope, TFA is clear about this:
"To get to principal, you need to put yourself on the critical path. To be effective as a principal and go beyond it, you need to actively remove yourself from it"
So whatever the sniper was doing that got him promoted to general, he must now stop doing it.
>> Those are choices. If you want to do that, you need a process that can support it.
__need__ is doing a lot of work here. There is no forcing function to get OEMs to do this ASAP: 1) the market doesn't really care that much 2) there are no regulations around this (and even if they were, can you immediately recall a tech exec going to jail for breaking the law ... )
This. Pixels are not more expensive than flagship Samsungs. If people cared and bought Pixels because they get the security updates, then Samsung (and the others) would follow. But people don't care, so the OEMs don't do it.
It's kinda weird to single out Samsung here, because they are pretty good with security updates and they explicitly talk about long security periods in their marketing. They are not as fast as Pixel, but somewhere mid-range and up (A5x) get monthly updates and they are usually 1-4 weeks behind Google.
It's the other vendors that are the issue. Even Fairphone is behind a lot (and they only release one model at a time).
The "(and others)" part was about including the other OEMs :-). I used the Samsung flagship as a specific example because it is very expensive, and people who buy it don't have the excuse of the price.
Yea all the “this bad thing will happen” discussion misses that the intent has been plain for at least a year now. The administration has plainly said what they would do during the election and they have rather faithfully executed on that plan. This isn’t about fixing things or saving people money it’s about doing what they want to do.
Its about inflicting as much pain in the american public as possible because if you cause indiscriminate damage its bound to also damage their enemies. Its a death cult.
Nah, these people are already rich beyond imagination, they want to be venerated, they have to see their enemies crushed. Money isn't enough for them, they want praise and the peasants haven't praised and honored them enough.
“I know this will be removed” is a pretty transparent attempt to sell the real reason for downvotes (you’re talking utter nonsense) as a made up one that suits your personal narrative.
Don’t be under the mistaken impression that you’re being smart here. It’s very transparent.
There is a long tradition in India, which started with oral transmission of the Vedas, of parallel cognition. It is almost an art form or a mental sport - https://en.wikipedia.org/wiki/Avadhanam
It is the exploration and enumeration of the possible rhythms that led to the discovery of Fibonacci sequence and binary representation in around 200 BC.
Sounds very much sequential, even if very difficult:
> The performer's first reply is not an entire poem. Rather, the poem is created one line at a time. The first questioner speaks and the performer replies with one line. The second questioner then speaks and the performer replies with the previous first line and then a new line. The third questioner then speaks and performer gives his previous first and second lines and a new line and so on. That is, each questioner demands a new task or restriction, the previous tasks, the previous lines of the poem, and a new line.
Evidence: see the amount of nonsense in the postgres rewritten in Rust story - https://news.ycombinator.com/item?id=48841676 where there are two contradictory claims made:
"postgres is so stable I will never trust a rewrite."
"covering 100% of postgres regression suite doesn't guarantee you have replicated every behavior."
Folks who are okay with LLMs think the regression test suite is the spec and is the guarantee of stability. How else can it be? If you are depending on some behavior not covered in the regression tests, how do you know the next minor release won't break you?
Folks who are against seem to imagine a platonic ideal of PG which conveniently is the original PG implementation by tautological definition. So no rewrite can ever meet their bar.