Tagged “rant”

What the heck is going on over at Proton?

tl;dr I’m no longer recommending Proton services to anyone and have moved off to other services.

Right, so up until recently I was a happy Proton customer at the annual unlimited tier. They were missing some convenience features and the integration of a document editor into the Drive product was pretty great.

That is, until they decided to shove an LLM into their core product and upend their security model and break some trust by the way they rolled it out. Let’s take a quick look at their timeline / speedrun of adding LLM features (and Crypto wallet, but we’ll get there in a sec).

  1. June 5th, 2024 - Proton releases the results of their 2024 community survey
  2. June 17th, 2024 - Proton announces transitioning to a non-profit structure
  3. July 8th, 2024 - Eamonn Maguire posts about building "privacy protecting AI" - Signaling that they were "thinking about" the problem (but clearly had already built the thing).
  4. July 18th, 2024 - Proton releases Proton Scribe - this is the LLM product.
  5. July 24th, 2024 - Proton releases a Bitcoin Wallet

Now that we’ve established the timeline, let’s break down the path from Survey to “releasing products that no one asked for”.

So I took this survey, and I’m now kicking myself for not screenshotting the questions, because I have a lot to say about shitty survey design.

I’m going to do some conjecture and speculation here, but I think that the survey was designed when they already had the Scribe product already well into development and that the questions were designed to make it seem like the product they were already developing was by popular demand. Like, you don’t develop a “privacy preserving” (it’s not) LLM in 1.5 months. They were already building this thing.

Let’s talk about those results. They provide this graph in their survey results post:

Chart showing that 29% of users want a "writing assistant"

My recollection of this question is that it was a multiple-choice, but not stack-ranked question. I’m not sure if that’s how it actually was, just how I remember it.

Regardless, I want to point out 2 things:

  1. The LLM answer in there doesn’t actually mention an LLM. It mentions a “writing assistant”. There are tons of things that do not use Large Language Models to do writing assistants, like checking grammar and spelling. The way they worded that answer was extremely misleadging.
  2. Only 29% of respondents said they wanted it.

So let’s move on to the second point they try to make in this survey that absolutely does not say what they say it does.

Chart showing percentage of users who "have used" AI

And here’s their analysis of that data:

Generative AI is one of the most significant developments in recent history, and it is supposed to lead to incredible gains in productivity. As more and more AI assistants come online, we asked the Proton community what they thought of these tools. Around 42% of respondents use an AI service regularly (at least once a month), and another 18% have never tried AI but are interested in it.

I just. sigh.

  1. The survey asked do you use Generative AI. It doesn’t ask why. It doesn’t ask if you find it useful. It doesn’t ask if you WANT IT IN YOUR FUCKING EMAIL CLIENT. This tells you nothing!
  2. “At least once a month” is not very much Generative AI usage.
  3. That’s still under half of your user base!

This question is poorly designed. I can’t tell if it’s poorly designed as an excuse to interpret the results the way Proton apparently wanted to, or if it was a genuine “surveys are hard to design” take, but asking “do you use Generative AI” without also asking “Do you use Generative AI for spicy roleplay” and “do you want us to put an LLM into your Email client” is ridiculous.

So, I don’t think this survey is a killer “Our users want this” result.

Testing the Waters - "What if we made a privacy focused LLM?"

Permalink to “Testing the Waters - "What if we made a privacy focused LLM?"”

On July 8th, about a month after producing the survey results, Eamonn Maguire posts about building "privacy protecting AI". This post reads to me like a love letter to Generative AI and was super suspicious to me. I posted on Mastodon in response to this post with the assertion that I'd immediately move off of Proton if they went forward with this.

Introducing an LLM, for our “business users”

Permalink to “Introducing an LLM, for our “business users””

10 days later, Proton releases Proton Scribe. Since you can't implement something complicated like this in 10 days... they already had this ready to go.

In their announcement post they make the following assertion:

In our 2024 community survey, more than 75% of Proton’s business users said they are interested in generative AI tools, but most were also concerned about a lack of data protections. Scribe was designed to be a secure alternative.

Right. Okay, I don’t think that’s what those survey results actually say (since you didn’t ask “do you want a secure alternative if we build one”).

When the backlash to this feature started over on Mastodon, they had the following response:

Screenshot of a Mastodon response

This excuse is flatly ridiculous. “We thought ‘hey they’re gonna do it anyway, let’s do it so they can be safe’”, is what that amounts to. That’s not a smart way to roll out features - and I don’t actually buy it.

Meanwhile, over on a Pivot to AI blog post on this subject, Amy and David point out the following:

Proton’s descriptions of Scribe are vague and waffly about their threat model. Your prompt — that is, the email you’re writing — is kept in plain text on their server, unlike emails you’ve sent or received, which are secure at rest. Proton promises they don’t log the prompts — but services like Apple, which many Proton users were trying to get away from, make only the same level of promise.

Proton then goes on to point out that the Model can run locally…. but only on Chrome, and only on systems with high enough system specifications. They went on to respond to this post with the following rebuttal on Mastodon:

Screenshot in response to a Mastodon post

I want to break down some of these things.

The feature is not “opt-in” if you’re on affected plans. It shows you this dialog:

Proton Scribe showing the opt-out dialog

Notice how that dialog doesn’t have a “no I don’t want this” option. To turn it off, you have to go into the settings. That’s an opt-out feature. Not opt-in.

The second thing is that there’s no such thing as an “Open source” model. The weights may be open, but Mistral has been super cagey about where they got the training data (also known as, they got it the same place all the other companies did - the internet, without permission).

Finally, if you or someone you’re chatting with decides to use the “zero logs server” the text is sent to them in the clear…. which I think very much does break their zero-knowledge model.

You’ve totally broken your privacy model for your users, willfully. For an LLM.

Ugh.

But that’s not the most ridiculous thing about this whole saga.

Literally six days later, Proton announces they’re launching a Bitcoin wallet of all things!

Pivot to AI covers this too and points this out:

If Proton was taking privacy seriously they’d have used Monero or Zcash — two cryptos that use zero-knowledge proofs to make transactions untraceable through the blockchain data trail. At least these have a use case, even if it’s buying Russian research chemicals off the darknets.

Yeah, exactly that. Bitcoin isn’t private, and no amount of you providing a non-custodial wallet is going to change that.

Here’s the security model. This is a reasonable attempt at the functionality, but it doesn’t make the basic idea any less dumb. Also, they recommend you use a bitcoin mixer so the Feds can bust you for money laundering too when they match you to your bitcoin address.

… That’s bad.

I really wish I knew. The timeline makes it clear to me that they had several potentially very unpopular features already in development and then timed the release with their developer survey. I suspect they were hoping to spin the results as a “you asked for it and guess what we delivered in record time!”.

The spaces I hang around in are very skeptical of both Generative AI and Crypto, and I think that Proton’s core individual customer is too. It feels like they’re speedrunning a “alienate our core customer base” but I honestly don’t know what they’re doing here.

It’s possible that their business users really do want this thing and have been asking for it for a long time. It’s also possible that the people running the show are adding features for a future investment or a sale. That doesn’t really jive with moving to a non-profit.

It’s also possible that they’ve always loved LLMs and Crypto and just happy to shove it into everything.

But, it doesn’t functionally matter for me. They’ve broken my trust, and that was the primary currency and reason I’d be willing to shovel money towards them. That takes a while to build. It takes an instant to destroy.

I think about that a lot. Oh well.

LLMs are a Cognitohazard

Cognitohazard image

So, I've been watching a bunch of folks whom otherwise have reasonable takes about LLMs who, under very specific circumstances, have extremely weird takes about how much an LLM is helping them do a thing, and their internal justifications for them. This isn't about anyone specific; for each of these points, I've seen variations of these from more than one person in every case. But they're happening enough that I want to talk about them. Don't just take it from me, this post from @xgranade makes the point better than I can:

It is amazing how many people I see claiming that AI has one good use. No one seems to agree on what that one good use is, and there's always a hell of a lot of goalpost shifting and special pleading involved.
For some reason a lot of folks quite reasonably follow the arguments against AI, but then partition off the one thing as immune or exempt from having to worry about any of the ethical and practical problems. - @xgranade@wandering.shop

Let me start, though. I think each and every one of these things ignores the externalities of running an LLM: The wholesale scraping of data frequently without consent (and the associated pushing of costs onto the site being scraped), the massive build-out of data centers that will never be used, a global chip shortage because the AI industry can just do that, the proliferation of slop and disinformation, and the trend of people experiencing psychosis when interacting with chatbots.

Okay, with that out of the way, let's dig in.

These are paraphrased. They are not direct quotes. If I use quotes, they are scare quotes to separate them from the rest of the sentence.

LLMs are great for Prototypes - I'd never have the time or ability to do this. It'd take me forever, or I'd never get it done.

Permalink to “LLMs are great for Prototypes - I'd never have the time or ability to do this. It'd take me forever, or I'd never get it done.”

Okay, so this one I've seen a whole lot lately. Namely, the claim that busy engineers are able to rapidly prototype something and get real time feedback on the prototype. This is usually associated with a sense that the engineer in question is "not a {insert type of} engineer" and so they'd never be able to complete such a task.

There are a bunch of problems with this approach, but I want to start with the claim that these folks would never be able to do something. That's patently untrue. If you're a software developer you absolutely have the capability to learn how something works and build out a prototype. The LLM's statistically average output is going to produce code that is prevalent. The most prevalent front end code being fed through these things is React. React has some very silly contrivances in it, but it's not an insurmountable mountain to learn the basics of React to roll out a prototype.

For code that is more difficult to learn, an LLM is going to have a harder time than you would, because the amount of samples in its training set are probably going to be a lot smaller.

The second problem with this attitude is that if you acknowledge that the code an LLM generates is substandard, and you want to treat it as throwaway.... it is never throwaway. Prototypes tend to make it, unmodified, into production. And if your usual engineers are busy, no one is going to substantively review this code.

Another issue is that the human involved in this process is not going to learn much from this process. Having the LLM write the code, and then try to make sense of the generated code is a failing proposition. The LLM does not have cognition. It produces code that is statistically likely. That's going to be a twisted mirror of the code it was trained on. Will it execute? Maybe! It might even produce code that executes on the first go. That's how slot machines work.

Also, this is why scaffolds and quickstarts exist?

I have the LLM do initial code review for me and point out things that are wrong

Permalink to “I have the LLM do initial code review for me and point out things that are wrong”

This one is the least offensive to me on this list, but it's still super annoying that I have to, for example, review a bunch of suggestions that Copilot makes on a GitHub PR that are either subtly wrong, or correct enough that accepting its suggestion won't break anything but it's also not doing anything of value (for example, making a comment more verbose).

I'm not convinced that this actually saves more time than it wastes. I've not had it find a single bug that would have caused a major problem had it made it into production, and it has (had its suggestions made it through) introduced bugs that would cause problems.

Perhaps I'm also being a bit petty, I can discuss a problem with a human and get to the point where we understand where there's a misunderstanding on either side - the LLM doesn't do that, it'll just agree with whatever you type to it, whether or not you're actually wrong. Because it's statistics.

It's worth pointing out that you, writing good unit tests and utilizing static code analysis (linters, anyone?) should be able to catch the same amount of stuff that an LLM would, deterministically.

Oh gods. Okay. So. I'm also a writer (fairly novice, granted), and I've found that there's nothing about typing at an LLM that actually generates quality ideas. It just spits out mild rehashes of old ideas and frequently the same variant of an idea a bunch of times in a row.

Why would I use electricity for this? I can grab a coffee, stare out the window, or talk to an inanimate object on my desk and generate a ton of stupid ideas that I can then iterate on, and I've found that works out way better than running a GPU to spit out "ideas".

I encourage you to give this a try. Find somewhere quiet, stare out of the window, and let your mind wander. Ideas will come.

LLMs are awesome at summarizing! I'm a busy professional and there's a lot of stuff for me to read!

Permalink to “LLMs are awesome at summarizing! I'm a busy professional and there's a lot of stuff for me to read!”

Nope. LLMs do not summarize, they shorten. They also frequently introduce inaccuracies or make up quotes when they're doing it. Often it'll also remove critical information, or emphasize the wrong thing. Overall, you cannot trust even the most advanced LLMs to accurately summarize a longer piece of text for you.

So what do you have to do if you want to be sure you have an accurate understanding of a longer text? You have to read it yourself. Which defeats the purpose of having the LLM shorten the text for you in the first place.

If you don't care about the accuracy of the summary...why are you reading it?

LLMs are great as a generalized search engine!

Permalink to “LLMs are great as a generalized search engine!”

So, this one is weird because it kinda depends on what you mean by "search engine". An LLM by itself isn't a very good search engine. It's maybe an okay reverse thesaurus or dictionary, but if you're looking for something with up-to-date information you're probably taking about something like Perplexity or the like.

Note, you can ask an LLM to generate "citations" and "link to your source" but it's reasonably likely to hallucinate those.

In the case of Perplexity and Google LLM search, what's actually happening behind the hood is a bit...insidious. They're using Retrieval Augmented Generation, namely, doing a web search first, and then feeding the first couple of results into the LLM to have it "summarize" (see previous comment) the results, and using a couple of techniques to inject the source reference into the "summary". Couple of major problems with this:

  1. It's only going to use the first couple of results for its summary, lest it risk overrunning its context window.
  2. It removes the results from their original context, so you can't be sure the summary doesn't include a shitpost from Reddit.
  3. "Hallucinations" still can crop up in those summaries.

Is this better than Google search has become? Maybe! Google has gotten pretty dang bad. I've resorted to running SearXNG on my own infra to get less bad results and no summaries.

Relatedly, we're seeing more and more hallucinated citations appearing in places they shouldn't so, uh.

This is just placeholder dialog, it'll be removed before the final product is out

Permalink to “This is just placeholder dialog, it'll be removed before the final product is out”

This one is a two-parter. Namely, the using of an LLM for "filler" text, and the use of a text-to-speech model to have verbal dialog - most commonly in games.

The first is a problem for even if you intend to remove it, bogus "filler" text will blend into the background. It will leave its mark on the final product. The odds that you'll remember to remove it decrease, when you could just leave the prompt in ("Emotional scene where Santa admits he made a pact with a devil for immortality", e.g.) and it'd be a lot clearer what needs to be rewritten. By not doing that, you're creating background radiation. Mediocre stuff that you may just miss when you're doing your next dialog pass.

Like a lot of the use cases I see for generative AI, it boils down to "just send me your prompt" rather than having the token spewer spit out tokens.

The latter is just... Like. If you need filler audio so you can "hear" the scene, just have someone around you read it out loud? Record it yourself? The same thing as before applies, if you don't intend to keep it in, there are other ways to get that out that do not require much more effort. You'll get more emotional nuance from a person who understands the scene, rather than the flat garbage that Amazon tried to pull.


That's all I've got energy for at the moment, but there are definitely more of these floating around out there, so I'll probably update this post later to add more, or make a follow-up. Haven't decided. I'm just so tired.