Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

How do you square away the idea that you do science for the increase in knowledge of human kind, but then say that a particular use of that knowledge is verboten?

I get the copyright aspect of this and I'm not arguing that here. I'm more asking about the moral / ethical idea of choosing who can benefit from your science.

Obviously there are the moral / ethical arguments about AI in general here to weigh against - those have been hashed out significantly elsewhere, and I'm not interested in debating them. What I'm asking about here is the impact on science by sharing it with tooling that distributes it in ways not generally considered when originally written.

A quick check of your post history suggests the frame that you work in strongly is privacy related research (observation - may be wrong). I'm curious how that impacts what you wrote here generally.

(Just to be perfectly clear, I'm not arguing your points here, trying to understand them better)



How are you sure that what big ai is doing is “the increase in knowledge of human kind”?

It is good for their business model; but it may lead to monopoly unlike anything acm ever had and very dubious prospect for academia.


What monopoly? What I'm seeing over these last 3 years is the fastest and fiercest competition I could imagine.


"it may lead to" is a future tense possibility.


I appreciate the grammar refresher, but my comment was about a monopoly on AI being very unlikely rather than it being an ontological impossibility.

What am I missing? Why should we be concerned about AI leading to a monopoly? And for which company?


I understand the LLM production companies have a funding mechanism but do they really have a business model?


Yes they do. They're drawing so much attention for the funding precisely because the business model is so attractive and has such a large TAM.


> How are you sure that what big ai is doing is “the increase in knowledge of human kind”?

The OP didn't say that.


Good question. Unfortunately, academic knowledge is widely ‘verboten’ already. Everything under paywall, researchers having to pay up to $10,000 to publish in open access in some venues, rare books unavailable even to top universities. Access to knowledge and information is increasingly difficult for everyone.

That said, what matters here is the social contract, what do I bring to society and what do we get from tech companies. For most people around the world, access to the typical leading models is out of reach. Not many on this planet can pay the subscriptions (or even API keys) that offer access to the best models. So I'm not buying the argument that tech companies are broadening access. What we're creating is a increasingly discriminatory society where the few get access to information, and the many don't.


Thanks for the reply here.

I guess I have some perspectives on a bunch of this. I'm for open sharing of academic work for all (but I'm not an academic, so my perspective is a consumer), so inferring your perspective here I think we agree on that. I maintain many open source (MIT/Apache2 licensed) libraries, and I've also worked in big tech (Amazon, OpenAI). I believe both in the idea of collective commons but also in the ideas that there should be the ability of people to sell software. There's tension in that social contract similarly, and it gets more complex when you look at copyleft.

I guess I'd be disappointed if this was just allowing big labs access and not more broadly allowing access to the ACM library for smaller open source models. Very much in agreement with your last points there.


Anything short of free access to the acm would ensure I'll fight ti burn down the acm instead. Not that they ever asked or cared what their members think. The organization can either choose to side with humanity or against it


> most people around the world, access to the typical leading models is out of reach. Not many on this planet can pay the subscriptions (or even API keys) that offer access to the best models. So I'm not buying the argument that tech companies are broadening access. What we're creating is a increasingly discriminatory society where the few get access to information, and the many don't

I’m sorry, but I can’t buy this argument. Making information more available does not make it more discriminatory. Nobody is saying it will only be available in the best models and withheld from other models or services like the ChatGPT free plan. Nobody is saying we’re going to make the original content inaccessible through the previous means after the LLMs are trained on it. Nothing about this shrinks access or makes it more discriminatory.

I understand that you’re upset about the use of the content, but I think you need to admit that your stance is the one trying to restrain use of the content. Training LLMs on it can only bring knowledge to a wider audience, not restrict it.

Whether or not that’s a good or fair idea is a separate discussion, but arguing that this makes access to the knowledge more discriminatory and locked away is 180 degrees backwards.


> Nobody is saying we’re going to make the original content inaccessible through the previous means after the LLMs are trained on it.

Is that not why they’re shredding the books when they’re done with them?


How is this a fair and reasonable assessment?

From here:

https://x.com/HedgieMarkets/status/2081534588485296565

A quote:

A federal judge ruled the practice is fair use because eliminating the original means only one copy exists at a time.

So no, the reason isn't to keep that information, even from competitors. It's about a legal ruling, which allows them to scan said books without infringing copyright. Judicial decisions and case law have lead to this outcome.

Please stop spreading rumours without any validity.


They asked a question. I saw no rumor spread.

Gotta love that ridiculous case law system :D


Well, sometimes a ? is rhetorical, and meant as a statement and not a question. But fair enough. And yes, I do believe the outcome is silly here.


First, this has nothing to do with the topic. ACM provides digital access. They don't have a single copy of a physical book which they're going to send to another company for destruction. Like I said, they're not deleting the source material or removing it from circulation.

As the other commenter posted, the original report that AI companies were shredding books was based on a second-hand retelling of a rumor, embellished for "AI bad" headlines.

There are actual bookshops talking about this are saying that most of the books are things like "How to master Microsoft Word 96" and that's why they're rare. They're also saying that the destination shipment is going to FBA (fulfilled by Amazon). They think it's an flipping operation trying to find arbitrage opportunities.

Also the reason AI companies have to destroy books is because they've been legally forbidden from using digital copies available. They had to pay a large settlement for it. So it's not some conspiracy to deprive the world of knowledge. It's what the courts told them they had to do.


Why do you think big AI is going to just share all it’s ‘knowledge’?


Because you can access it for free? As it was the case from the first day LLMs became a thing in public consciousness?

Or did I miss the change, and ChatGPT, Claude and Gemini no longer have a free tier anymore? And then the next tier that costs peanuts for anyone in the west, that gives you more access to slightly fresher models?


That is interacting with it - which people are getting charged more and more for.

The first ones are free.

And as long as they keep the weights proprietary (which all the major players are for their primary models), they can decide to charge whatever they want later.


> which people are getting charged more and more for.

Completely false. The price for LLM inference at a given level of model intelligence has been dropping like a rock.

There are more expensive models available, but you don't have to use them. The same LLM knowledge that was available a couple years ago is now free to download and run on your laptop.


Says no one managing a corp budget!


I’m talking about price for a given output.

You’re talking about companies using more tokens.

Completely different concepts. As I said, the price for a given quality of LLM output continues to decline. Has nothing to do with companies using more tokens.


I’m talking about they can charge you whatever, and you don’t own anything so you don’t get a say.


Whatever was released in the open as weights is forever out there and beyond reach of any corporate interest.


it’s like you can’t even read


Exactly. Companies like Anthropic do not want open access. See their FUD against "distillation attacks"


I keep coming to this thought too. If we really get to the point that AI can produce all the software we need, do any knowledge work we need, why wouldn't Big AI just use it to do that and take you out of the loop? Are "AI developers" making the flashlight app in the App Store right now?


> but then say that a particular use of that knowledge is verboten?

https://en.wikipedia.org/wiki/Paradox_of_tolerance

Without material values, you will be lost and confused.

Granted, the ACM is hardly a bastion of anything but self interest and greed.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: