Tagged “blog”

RSS feed added to blog

So thanks to the RSS Plugin for 11ty I now have these posts available via RSS! The feed is linked in the footer and is available at /feed.xml.

I intend to do the same thing to the reviews page just as soon as I finish up my headphone reviews page. Which I'll definitely do, one of these days....

Alternative Video Hosting

I've fallen down a super deep Fediverse rabbit hole. It feels kinda like for the first time I can play around with stuff and not have my Extra Life videos just living on Twitch or Youtube. As a result, I've put up my own PeerTube instance up on vid.cthos.dev and am slowly moving the Extra Life videos there. They're going to stay up on Youtube as well, but here's a much better place to move things. Going to see if I can also make copies of my old talks, but odds are good I don't have the rights to reproduce them. In any case, here's a copy of my Sucker for Love run!

Steam Deck vs Ayaneo 2

tl;dr - If you're most people, get a Steam Deck. If you've got similar use cases to me Ayaneo 2 might be for you.

I don't think I've ever posted about this before, but I wanted to take a bit of time to talk about one of my pretty-obvious hobbies. Despite having way too many side projects, and working too many hours in the week, I still like to try and find some time to play video games.

After working 50-60 hours a week, I like to unwind - but I don't really have a great place to play console games - my PS 5 is at the same desk where I work, and at the end of the day I really like to get away from all of that. I solved this problem with a gaming laptop (13" Razer Blade - love that thing), but it still required a fair amount of effort - keeping a controller near it, having a lap desk so it didn't burn my legs, etc.

Enter the Steam Deck. I preordered one the moment it became available (well, like an hour in because the servers were crashing) and then waited well over a year for it to arrive. But when it arrived.... oh boy was that thing awesome. The controller layout is top-notch, it's comfortable to hold (if a bit heavy), and since I was playing on the couch - the laughably short battery life wasn't an issue. It solved basically all of my problems... except for one major one.

You see, dear readers, every year I do Extra Life - and every year I need to stream for a full 24 hours. While concievably I could do this on the Steam Deck - the amount of CPU things like OBS take up would be pushing the Deck's hardware to the limit. That, coupled with the screen being a bit sub-par in terms of color saturation and resolution, I started to look for alternatives.

Specifically alternatives that support an eGPU.

Enter the Ayaneo 2. Now, for those of you who don't know about Ayaneo - they're a small Chinese company who (for a good while) had a reputation for the best support of the not-Valve handhelds (because no one is going to match Valve in this space - they're simply the largest player), with a reputation for quality. So, I figured I'd give the 2's Indiegogo campaign a go1.

Long story short, they had a number of issues with that campaign and a lot of folks were having issues with either delivery, faulty hardware, support, or the like. I was not one of those people - I had zero issues with the 2 when it arrived. Let me tell you, I love this thing. The screen is beautiful, it's comfortable to hold (though the Deck has the superior inputs), and after a lot of tinkering and driver updates - I got it to work with my eGPU. Problem solved, I will be able to stream this year without resorting to hijacking what is now my Fiancee's gaming laptop and have a solution for when I need to play those games that have a bit higher graphical demands.

Deck Battle, iPad for scale

So, how do the two consoles stack up to each other?

  • Input: The back buttons are amazing and really really useful. Especially when you remap one of them to "right stick depress" - makes toggling run in a bunch of games much easier. Also, trackpads are exceptionally useful for certain games.
  • Comfort: It's just slightly nicer to hold for longer periods due to where the sticks are placed vs. the d-pad.
  • Software: The Software experience on SteamOS is really tight. It just workstm. That cannot be said about the Win 11 experience on Ayaneo - and Ayaspace is semi-necessary for changing TDP and mouse support, but it's clunky.
  • Screen: The screen is beautiful, and higher resolution. Even when dropping the resolution on AAA games to get better performance, it just looks better.
  • Performance: You can set the TDP on this thing fairly high, higher than the Deck, and at higher TDPs you get better performance for your trouble.
  • eGPU support: The Deck doesn't have a USB-4 port so it cannot do eGPUs even if it wanted to, so.... here we are.

All in all, they're both great hardware and the Ayaneo 2 is working great for me - but for most people, I think the Steam Deck is going to be the way to go -


1: Full disclosure, I'd already gotten an Air from them at this point, and enjoyed it. The air is my travel console.

Reworking the Blog

So! The time has come once again to rework the blog. I'll be doing a lot more writing in a couple of different places.

First off, for what I'll be working on professionally, head on over to this Cthonic Studios Post where I outline several of the TTRPG and Software projects I'll be doing in 2024 (I'm super excited about all of this).

Cthonicstudios.com is where I'll be doing all of my "professional" blogging, complete with musings on all the philosophy of TTRPG stuff I've been thinking about, as well as updates on the games I'll be making.

Alextheward.com by contrast will be the place where I do a lot of musing about tech, personal updates, and the like. I'll likely link back and forth from one site to another - but one thing's for sure, I'll be doing a lot more blogging on both sites in the very near future.

Likewise, I'll be setting up some mailing lists - if you'd like to get updates from me it's relatively easy to subscribe by creating an account on Cthonicstudios, but there will also be a separate mailing list for here.

Hope you all stay tuned for what's next, because I'm going to do my darndest to make 2024 a fun and transformative year.

P.S. I may decide to migrate this site over to either a different host (away from Netlify), and I may also move it from 11ty to Ghost (but I also love 11ty - but I enjoy the editorial experience of Ghost and the nice features it enables). So we'll see. For now, I rest.

Passkeys: your friendly password replacement

Update: 04-26-24 - whooboy, so apparently I've missed some marketing nonsense around Passkeys. At some point, someone decided that Passkey === Resident Key and that apparently stuck? I wasn't aware of that, but I'm gonna link to a technical blog here where you can read more about the chaos.
As a result, this means the information below is a little...weird? Like, it's all still accurate except apparently "passkey" means "usernameless / discoverable key" in public parlance. Which is dumb? I don't like that. Yubikeys can only store up to 25 resident keys which makes them way less useful. Very handy in say, the Apple ecosystem still, but hey. I've added another sidebar in the context.

One of my absolute favorite topics to rant about continually is identity and access management (specifically authentication). One of my friends prompted me to write about this after needing some questions about passkeys answered.

Anyhow, one of the most thrilling innovations in identity managements in recent years that is finally seeing wider adoption is Passkeys. Put simply, Passkeys replace the need for passwords by shunting the authentication checks to a combination of something you have (a phone, or security key) with something you are (biometrics) to authenticate to a website or a service. No longer do you need to use a password manager to log in securely to a site, or remember a secret phrase, so long as you have your device and you remain you - you can access your content.

They’re unphishable, because you can’t just hand them over - you physically need to be in possession of the device. You can’t forget them. There’s nothing to remember. Stealing an entire database of public keys gets you basically nothing, you can’t really use those public keys to impersonate anyone - and because they’re website-specific, they’re immune to credential stuffing. They’re (usually) behind biometrics - so even if your device is stolen, the thief also must worry about faking your face or fingerprint. It makes the cost / time investment in stealing credentials much harder.

The way it does this is all behind-the-scenes and the end user doesn’t need to really understand how any of it works, but today we’re going to dive into that just a bit so you can peer behind the curtain and understand the limitations and risks of bad implementations.

So how does this passkey thing work anyhow?

Permalink to “So how does this passkey thing work anyhow?”

Easy! Public Key Cryptography! … wait that’s not easy? Right, okay, let’s start from the beginning. Passkeys are based on the concept of public key cryptography, in which a system generates a pair of keys on a device - a public key and a private key. The public key can encrypt data so that only the private key can decrypt it, and vice versa. Because the private key never leaves the device, you can hand out the public key so that folks can encrypt data so that only you (the device that holds the private key) can decrypt it. You can also create signatures which simply prove possession of a key without encrypting the data.

Passkeys, using a protocol called WebAuthN, take advantage of this by exchanging public keys back and forth so that when you register / log in, that you have your private key becomes your authentication proof.

The registration flow for a “brand new account” essentially works like this (I’ll be leaving out a fair amount of detail for the sake of understanding - if you want the full picture, the WebAuthN spec, Google, and Apple all have wonderfully detailed guides for implementing Relying Parties).

Okay, so for this section a "resident key" can also be usernameless, meaning the app can just straight sign you in without asking for an identifier - and apparently the wider passkey marketing means that's necessary. Which it's not actually necessary, but here we are.

  1. You visit a website and click “register”.
  2. For non-resident keys only: The website prompts you for some sort of identifier (most commonly email address, but could be username or anything, really) - or generates one for you.
  3. Your device prompts you to create a new passkey, and has you biometrically authenticate (face/touch ID, on iOS).
  4. For Resident Keys: Your device generates a new public / private key pair specific to that website’s domain name and sends the public key to the website. For Non-Resident Keys: a seed gets generated and stored against the identifier - that seed is what is used by the hardware token in conjunction with the private key resident on the token to ber unique per-site (still specific to the website's domain).
  5. The website stores your public key along with your identifier for future use; you are now registered.

This flow can also work for an existing legacy username/password account to add a new passkey to that old account - the only difference is the website already knows what account to associate the passkey to.

For login, the operation is very similar, but there’s an additional challenge mechanism in place.

  1. You visit the website, enter your username/email/whatever to identify which account to log into (the browser can also store this for you, so it knows you already).
  2. The website issues a challenge back to your device, basically: “I need you to prove you have this private key by using it to sign back this data I’m sending you”.
  3. Your device prompts you to use the stored passkey for this website and has you biometrically authenticate.
  4. Your device creates an attestation - which is the response to that challenge, but signed using the private key.
  5. The website uses the public key (or seed as necessary) to check that the assertion was signed correctly. If it was, then you are now logged in! If it wasn’t, the site rejects your credentials.

If you’re interested in the technical details, but do not want to read the spec, I recommend this guide on WebAuthN.guide, or the Google implementation Sandbox.

How is this different than using FaceID in an App?

Permalink to “How is this different than using FaceID in an App?”

It’s really not all that different - many of the same mechanisms that Apple uses for in-app FaceID also apply to Passkeys, they’re built on very similar tech.

Same deal for Google’s biometric implementations, a lot of the same plumbing powers the overall ecosystem.

Cool, so uh, what happens if I drop my phone in the lake?

Permalink to “Cool, so uh, what happens if I drop my phone in the lake?”

You may have noticed that there’s a potential problem if the identity is tied to a single device - how do you log in if you’ve lost the device? There are a couple of ways to handle this on both the device side and the provider / website side.

Apple and Google handle this by synching that private key to your cloud account on either platform - encrypted by the master password on your account. That way, you can use the same passkey on any device you own in the same ecosystem, protected by the same biometrics. For 95% of people, this is sufficient. Even if you lose your phone, so long as you can still get into Google / iCloud your credentials are safe. This means you still need one password - for your cloud account, but you can also protect that via something like a Yubikey stored in a lockbox or something.

On the provider side, they can offer you the ability to enroll multiple passkeys, so that you can keep one as backup on a physical security key. They can also offer robust backup options - currently this is typically 10 or so random “break glass” codes that you can print off and keep in a safe somewhere. They also still have the option of offering traditional SMS recovery, but your security is only as good as the weakest recovery method - and SMS is extremely insecure.

For those extremely paranoid, you can buy hardware authenticators and use those to generate passkeys, but you are responsible for backups at that point - you want to keep one somewhere safe in the event you drop your token in a gorge.

Note: Yubikeys support at most 25 resident "usernameless" keys.

Is this really the future of authentication?

Permalink to “Is this really the future of authentication?”

I think so. To put it extremely bluntly, passwords are awful. They can be stolen. You can forget them. We’ve been trying for years to get the less tech-savvy among us to adopt (and pay for!) password managers to create strong random passwords so that perhaps we can prevent credential stuffing. That hasn’t worked, because the barrier to entry is high and confusing.

The barrier to entry for passkeys is basically zero: “Do you want to let FaceID / TouchID manage this login for you?”. For most people, it will “just work”, and for the rest of us, our password managers or hardware tokens are for us to manage.

I don’t know if passwords will ever fully go away, but I sure hope that passkeys take out a huge chunk of them, and soon.

Let's do some "Game Development"

Shutterstock vector art of some hands typing in front of a monitor showing an animation editor
Kit8.net @ Shutterstock #1498637465

The other day, a friend asked me a complicated question wrapped in what seems like a simple question: “How long would it take to make some simple mini-games that you could slap on a web page?”

My answer, as many seasoned developers will be familiar with (and of course it’s coming out of the architect’s mouth) was “It depends on the game and the developer. I could probably churn out something pretty good in a few days, but it’d take someone more junior longer”.

This set off the part of my brain that really wants to test out just how fast I could do a simple game in engines that I’m not familiar with. Thus, I decided to ignore my other responsibilities and do that instead! Mostly kidding about the responsibilities thing, I’m ahead of my writing goals for Nix Noctis, so I had a couple of hours to spare (and a free evening).

How long would it take to make a simple “Simon” game in GDevelop: a “no code” game engine? I wanted to start with GDevelop for a few reasons:

  1. I wanted to simulate the experience someone with little coding experience would encounter. Knowing that I’ve internalized some concepts that would make the experience easier (or, in some cases, harder) for me.
  2. You can run the editor in a web browser, and that’s bonkers.
  3. I want to “pretend” to be a beginner; only use basic and easily searchable features.
    1. I mean, I am a beginner at game dev, but I do understand several the concepts.

With those things in mind, I set out to remake Simon. Here were my design constraints:

  • Four Arrows, controllable by clicking them or via the keyboard.
  • Configurable number of “beeps” in the sequence
    • The game would not increase by one each time (though it easily could)
  • Three Difficulty levels which control the speed that the sequence is shown
  • Bonus points for sound effects

REMEMBER! I’m a beginner at GDevelop. You’re likely going to see something and say “hey that’s dumb, you should have done it another way”. Yes. Exactly.

The editor experience in GDevelop is actually really nice, especially when you can just get into it via clicking into a browser. I found adding elements to the page very intuitive. Sprites are simple, and adding animation states to them is effortless. Creating the overall UI took me probably 20 to 30 minutes to iteratively build out a structure I was happy with, it was fast. Another fun thing I discovered was that they have [JFXR] built into the editor, and that was a delight.

Screenshot of GDevelop's UI Editor

What was not so quick was wiring up the game logic to the elements on the page. I’ve looked at some GDevelop tutorials before, and if you’re treading a path that’s covered by one of their “Extensions”, you’re going to have a great time. A 2D platformer will be a breeze because you can simply attach a behavior to the sprites in your game and go. There are a bunch of tutorials on making a parallax background for really cool looking level design. Simple!

What is not so simple is if you fall outside those behaviors and need to start interacting with the “no code” editor. On one hand, the no code editor is nice! The events and conditionals are intuitive if you’re approaching it in certain ways. They even let you add JS directly if you know what you’re doing (though they recommend against it). On the other hand, I can see this getting quickly messy. In my limited experience with the engine, I could not find a good way to reuse code blocks. This will come up later.

Sidebar, dear reader, I believe this is where they would say “you should make a custom behavior to control this”. I’m not sure a beginner would think to do this, but I thought about it and said, “I’ll just duplicate the blocks”.

Screenshot of GDevelop's Code editor

As I worked through this process, I ran into a number of weird stumbling blocks that slowed down my progress while I tried different things.

Things that were surprisingly straightforward

Permalink to “Things that were surprisingly straightforward”

GDevelop has many layers of variables, Global, Scene, Instance, and so forth. Easy to understand, fairly easy to access and edit.

Were I making this game in pure JS, generating the sequence would be a pretty “simple” one-liner (It’s not that simple, but hey, spread operator + map is fun! I’d expand this in a real program to be easier to understand):

// Fill a new array of size number_of_beeps with a digit between 0 and 3 to represent arrow directions
let sequence = [...Array(number_of_beeps)].map(() => Math.floor(Math.random() * 4));

Generating the sequence was remarkably easy once I figured out how to do loops in GDevelop; they’re hidden in a right-click menu (or an interface item), but the “Repeat x Times” condition is precisely what we needed.

Screenshot of the Repeat x Times condition

Likewise, doing the animation of the arrows was pretty direct. All you need to do is change the animation of the arrow, use a “wait” command, and then turn it back. Easy!

Screenshot of GDevelop's editor

Turns out it’s not actually that easy. The engine (as near as I can tell) is using timeouts under the hood which means they’re not fully blocking execution of other tasks in sibling blocks while this is happening. Which means…

Wait for x Seconds is weird / doesn’t work right

Permalink to “Wait for x Seconds is weird / doesn’t work right”

Okay, so when you’re playing Simon, the device will beep at you, light up for a non-zero number of seconds, and then dim. It should do that for each light. If you’ve never experienced this wonder of my childhood, watch this explanatory video from a random store:

Now that we know how that works, we want to emulate that in GDevelop. The first thing I tried was to simply put the “Wait” block at the end of a For...In Loop. Yeah, remember what I said about timeouts? Those don’t block. The loop would just continue and totally ignore the wait. I think that’s a major pitfall for new devs, they’re not going to understand the nuance of how those wait commands function under the hood.

The second thing I tried is the “Repeat Every X Seconds” extension to do the same thing. I couldn’t get it to even fire, and I still don’t know why.

Anyhow, I settled on using a timer to do the dirty work. Here’s how our “Play the sequence loop” wound up looking at the end:

Screenshot of Play Sequence Loop

Conditionals / Keyboard Input Combined with Mouse Clicks

Permalink to “Conditionals / Keyboard Input Combined with Mouse Clicks”

There’s another thing I could not figure out. I wanted to have both mouse clicks and keyboard input control the “guess” portion of the code, so I sensibly (imo) attempted to combine those into a single conditional. That wound up being…very weird, and there’s still a bug related to it.

First off, there are AND and OR conditionals. It took me a bit to find them, but they do exist. So, with a single OR conditional and a nested AND conditional, I set out with this:

Screenshot of the Nested Conditionals

This mostly works. However, for whatever reason, if you use the keyboard to input your guess, the arrow animation does not play. I do not know why. It works if you click. I can prove the conditional triggers in either case. It’s just that the animation does. Not. Work. Maybe one day I’ll figure it out, but I chose not to.

Struggling past some of those hurdles, it took me about 5 hours to meet my original design goals in GDevelop. Not terrible for not knowing anything about GDevelop besides that it exists.

Here’s a video of the final product:

Okay! We’re done here. Post over.

Nah, you all know me, I couldn’t stop there.

After I’d done the initial experiment, I was curious if I could work faster in a game engine that has real scripting support.

The short answer is “yes”. It took me about 2.5 hours to complete the same task in Godot (with some visual discrepancies). I think the primary speed gain was from the fact that the actual game logic was much more intuitive to me and the ability to wire up “signals” from one element to the main script code made it much faster to do some of the tasks I was fighting in GDevelop.

Godot also has an await keyword which blocks execution like you’d expect it to, which is outstanding.

I did run into one major issue that I had to do a fair amount of research to solve:

AnimatedSprite2D “click” tracking is surprisingly difficult

Permalink to “AnimatedSprite2D “click” tracking is surprisingly difficult”

The only issue I had was that when I needed to determine if the user had clicked on an arrow, I had to jump through some interesting hoops to detect if the mouse was over the arrow’s bounding box.

While regular sprites have a helper function get_rect() which allow you to figure out its Rect2D dimensions, AnimatedSprite2Dvery much do not (you have to first dig into the animations property, and then grab the frame it’s currently on, and then you have to get its coordinates and make your own Rect2D. Gosh, I’d have loved a helper function there).

I think the expectation is you’d have a KinematicBody2D wrapping the element, but as the arrows are essentially UI, that didn’t make any sense to me. I’ll need to dig a bit further into how Godot expects you to build a UI to do all of that, but hey, I got it working relatively quickly.

Changing the text of everything in the scene was really bizarre due to how it’s abstracted via a “Theme” object that you attach to all the UI elements? Still haven’t quite figured that out. It was really easy in GDevelop. Not so much in Godot.

Yeah, so, I liked working in Godot more because it was easier to make the behaviors work, and I was getting exhausted at the clunkiness of the visual editor. Here’s the final product:

For me, working in both of these engines for fun was a positive experience and I can see myself using GDevelop for some quick prototyping, but personally, I like Godot’s approach to the actual scripting portions of the engine. Because I have a lot of software development experience, it’s much easier for me to just write a few lines of code over having to navigate the quirks of the interface.

I think GDevelop is perfectly serviceable, though. It looks like everything in the engine does have a JS equivalent, so you really could just write JS if you wanted to. If they exposed that more cleanly, I think it’d be pretty great for many 2D needs.

But I’m not a game dev, this is just me tinkering around and giving some impressions. Go try them out for yourself, they’re both easy to get started with!

Own Your Content

Shutterstock vector art of some computer magic
Andrey Suslov @ Shutterstock #1199480788

Not too long ago Anil Dash wrote a piece for Rolling Stone titled “The Internet Is About To Get Weird Again” and it’s been living rent free in my mind ever since I read it. As the weeks drag on, more and more content is being slurped up by the big tech companies in their ever-growing thirst for human created and curated content to feed generative AI systems. Automattic recently announced that they’re entering deals to pump content from Tumblr and Wordpress.com to OpenAI. Reddit, too, has entered into a deal with Google. If you want a startling list, go have a look at Vox’s article on the subject. Almost every single one of these deals is “opt-out” rather than “opt-in” because they are counting on people to not opt-out, and they know that the percentage of users who would opt-in without some sort of compensation is minimal.

Lest you think this is a rant about feeding the AI hype machine, it’s not (though you may get one of those soon enough). This is more of a lament from the last several decades of big social medial companies first convincing us that they are the best way to reach and maintain an audience (by being the intermediary) and then taking the content that countless creators have written for them and then disconnecting the creator from their audience.

Every bit of content you’ve created on these platforms, whether it’s a comment or a blog post for your friends or audience is being monetized without offering you anything in return (except the privilege of feeding the social media company, I guess). Even worse, getting your stuff back out of some of these platforms is becoming increasingly difficult. I’ve seen many of my communities move entirely to Discord, using their forums feature. However, unlike traditional website forums you cannot get your forums back out of Discord. There’s no way to backup or restore that content.

I’ve personally witnessed a community lose all of its backup content due to a leaked token and an upset spammer. It was tragic and I still mourn (but hey, we’re still there).

In one way, this is the culmination of monetizing views. As Ed Citron argues in Software has Eaten the Media, the trend from many of these social media companies has been “more views good, content doesn’t matter”. We’ve seen this show before, Google has been in a shadow war with SEO optimizers for over a decade, and they might have lost. The “pivot to video” Facebook pushed was a massive lie, and we collectively fell for it.

So what do we do about this? One thing I’m excited to see that Mr. Dash rightly points out is that there’s a renewed trend of being more directly connected to the folks that are consuming your content.

Own a blog! Link to it! Block AI bots from reading it if you’re so inclined. Use social media to link to it! Don’t write your screed directly on LinkedIn - Don’t give them that content. Own it. Do what you want to with it. Monetize it however you want, or not at all! Own it. Scott Hanselman said this well over a decade ago. Own it!

Recently, there was a Substack Exodus after they were caught gleefully profiting off of literal Nazis. Many folks decided to go to self-hosted Ghost instead of letting another company control the decision making. Molly White of Citation Needed (who does a lovely recap of Crypto nonsense) even wrote about how she did it. Wresting control away from centralized stacks and back to the web of the 90s is definitely my jam.

Speaking of Decentralization, we’ve also got Mastodon and Bluesky that have federation protocols (Bluesky just opened up AT to beta, which is pretty cool) allowing you to run your own single-user account instances but still interact with an audience (which is what I do).

Right, anyhow, this rant is brought to you by the hope that we’re standing on the edge of reclaiming some of what the weird web lost to social media companies of yore.

Edit: 03-10-2024: Turns out there's a name for this concept! POSSE: https://indieweb.org/POSSE. My internet pal Zach has a fun 1-minute intro on the concept: https://www.youtube.com/watch?v=X3SrZuH00GQ&t=835s. Go watch it!

How I’m approaching Generative AI

The Duality of AI
Lidiia Lohinova @ Shutterstock #2425460383

Also known as the “Plausible Sentence Generator” and “Art Approximator”

This post is only about Generative AI. There are plenty of other Machine Learning models, and some of them are really useful, we’re not talking about those today.

I feel like every single day I see some new startup or post about how Generative AI is the future of everything and how we’re right on the cusp of Artificial General Intelligence (AGI) and soon everything from writing to art to music to making pop tarts will be controlled by this amazing new technology.

In other words, this is the biggest tech hype cycle I’ve personally witnessed. Blockchain, NFTs, and the like come close (remember companies adding “blockchain” to their products just to get investment in the last bubble?) and maybe the dotcom bubble, but I think this “AI” cycle is even bigger than them all.

There are a lot of reasons for that, which I’m going to get into as part of this … probably very lengthy post about Generative AI in general, where I find it useful, where I don’t, where my personal ethics land on the various elements of GenAI (and I’ll be sure to treat LLMs and Diffusion models differently). So, by the end of this, if I’ve done my job right, you’re going to understand a bit more about why I think there’s a lot of hype and not a lot of substance here — and how we’re going to do a lot of damage in the meantime.

Never you worry friends, I’m going to link to a lot of sources for this one.

If you’ve been living deep in a cave with no access to the news you might not have heard about Generative AI. If you are one of those people and are reading this, I envy you - please take me with you. I’m going to go ahead and define AI for the purposes of this article because the industry has gone and overloaded the term “AI” once again.

I’m going to be very constrained to “Generative AI”, also known as “GenAI”, of two categories: Large Language Models (LLMs) and Diffusion Models (like Dall-E and Stable Diffusion). The way they work is a little bit different, but the way they are used is similar. You give them a “prompt” and they give you some output. For the former, this is text and for the latter this is an image (or video, in the case of Sora). Sometimes we slap them together. Sometimes we slap them together 6 times.

Examples of LLMs: ChatGPT, Claude, Gemini (they might rename it again after this post goes live because Google gonna Google).

I’m going to take my best crack at summarizing how this works, but I’ll link to more in-depth resources at the end of the section. In its most basic terms, an LLM takes the prompt that you entered and then it uses statistical analysis to predict the next “token” in the sequence. So, if you give it the sentence “Cats are excellent”, the LLM might have correlated “hunters” as the next token in the sequence as statistically 60% likely. The word “pets” might be 20%. And so on. It’s essentially “autocomplete with a ton of data fed to it”.

Sidebar, a token is not necessarily a full word. It could be a “.”, it could be a syllable, it could be a suffix, and so on. But for the purposes of the example you can think of as words.

What the LLM does that makes it “magical” and able to generate “novel” text is that sometimes it won’t pick the statistically most likely next token. It’ll pick a different one (based on its Temperature, Top-P, and Bottom-P parameters), which then sends it down a different path (because the token chain is now different). This is what enables it to give you a Haiku about your grandma. It’s also what makes it generate “alternative facts”. Also known as “hallucinations”.

This is a feature.

You see, the LLM has no concept of what a “fact” is. It only “understands” statistical associations between the words that have been fed to it as part of its dataset. So, when it makes up court cases, or claims public figures have died when they’re very much still alive, this is what’s happening. OpenAI, Microsoft, and others are attempting to rein this in with various techniques (which I’ll cover later), but ultimately the “bullshit generation” is a core function of how an LLM works.

This is a problem if you want an LLM to be useful as a search engine, or in any domain that relies on factual information, because invariably it will make fictions up by design. Remember that, because it’s going to come up over and over again.

  1. Stephen Wolfram talks about how ChatGPT works.
  2. On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? 🦜

I don’t understand diffusion models as well as I understand language models, much like I understand the craft of writing more than I do art, so this is going to be a little “fuzzier”

Examples of Diffusion Models: Dall-E 3 (Bing Image Creator), Stable Diffusion, Midjourney

Basically, a Diffusion model is the answer to the question “what happens if you train a neural network on tagged images and then introduce progressively more random noise?” The process works (massively simplified) like this:

  1. The model is given an image labeled “cat”
  2. A bit of random noise (or static) is introduced into the image.
  3. Do Step 2 over and over again until the image is totally unrecognizable as a cat.
  4. Congrats! You now know how to make a “Cat” into random noise.

But the question then becomes “can we reverse the process?”. Turns out, yes, you can. To get a from a prompt of “Give me an image that looks like a cat” the diffusion model will do the process in essentially reverse:

  1. We generate an image that is nothing but random noise.
  2. The model uses its training data to “remove” that random noise, just a bit
  3. Repeat step 2 over and over again
  4. Finally, you have an image that looks something akin to a cat

Now, on this other side, your model might not have generated a great cat. It doesn’t know what a cat is. So, it asks another model: “Hey, is this an acceptable cat?” Said model will either say “nope, try again”, or it will respond with “heck yes! That’s a cat. Do more like that”.

This is Reinforcement Learning - this is going to come up again later.

So, at it’s most “basic” representation the things that are making “AI Art” are essentially random noise de-noiserators. Which, at a technical level is super cool! Who would have thought you could give a model random noise garbage and get a semi-coherent image out of the other end?

  1. Step by Step Visual Introduction to Diffusion Models by Kemal Erdem
  2. How Diffusion Models Work by Leo Isikdogan (video)

These things are energy efficient and cheap to run, right?

Permalink to “These things are energy efficient and cheap to run, right?”

I mean, it’s $20/mo for an OpenAI ChatGPT pro subscription, how expensive could it be?

My friends, this whole industry is propped up by a massive amount of speculative VC / Private Equity funding. OpenAI is nowhere near profitable. Their burn rate is enormous (partly due to server costs, but also training foundational models is expensive). Sam Altman is seeking $7 Trillion dollars for AI chips. Moore’s law is Dead, so we can’t count on the cost of compute getting ever smaller.

Let’s also talk about the environmental impact of some of these larger models. Training them requires a lot of water. Using them uses way less water (well, as much as running a power-hungry GPU would require), but the overall lifecycle of a GenAI large foundational model isn’t exactly sustainable in the world of impending climate crises.

One thing that’s also interesting is there are a number of smaller, useable-ish models that can run on commodity hardware. I’m going to talk about those later.

I think part of what’s fueling the hype here is only a few companies on the planet can currently field and develop these large foundational models, and no research institutions currently can. If you can roll out “AI” to every person who uses a computer, your potential and addressable markets are enormous.

Because there are only a few players in the space, they’re essentially doing what Amazon is an expert at: Subsidize the product to levels that are unsustainable (that $20/mo, for example) and then jack up the price later once you’ve got a captive market that has no choice in the matter anymore.

Go have a watch of John Stewart’s interview with FTC Chair Lina Khan, it’s a good one and touches on this near the end.

We’re already seeing them capture a lot of market here, too, because a ton of startups are building features which simply ask you, the audience, to provide an OpenAI API key. Or, they subsidize the API access cost to OpenAI through other subscription fees. Ultimately, a very small number of players under the hood control access and cost…. which is going to be very very “fun” for a lot of businesses later.

I do think OpenAI is chasing AGI…for some definition of AGI; but I don’t think it’s likely they’re going to get there with LLMs. I think they think that they’ll get there, but they’re now chasing profit. They’re incentivized to say they’ve got AGI even if they don’t.

Cool! So we modeled this on the human brain?

Permalink to “Cool! So we modeled this on the human brain?”

I’m getting pretty sick of hearing this one, because the concept of a computer Neural Network is pretty neat but every time someone says “and this is how a human brain works” it drives me a little bit closer to throwing my laptop in a river.

It’s not. Artificial Neural Networks (ANNs) were invented in the late 1960s, and were modeled after a portion of how we thought our brains might work at the time. Since then, we’ve made advances with things like Convolutional Neural Networks (CNNs) starting in the 1980s, and most recently Transformers (this is what ChatGPT uses). None of these ANN models actually model what the human brain is actually doing. We don’t actually understand how the human brain works in the first place, and the entire field of neuroscience is constantly making discoveries.

Did Transformer architecture stumble upon how the human brain works? Unlikely, but, hey, who knows. Let’s throw trillions of dollars at the problem until we get sentient clippy.

Look, I could get into a lengthy discussion about whether free will exists but I’m gonna spare you that one.

Wikipedia covers this better than I could, so go have a read on the history of ANNs.

AI will not just keep getting better the more data we put in

Permalink to “AI will not just keep getting better the more data we put in”

Couple of things here: it’s really hard to model how well a generative AI tool is doing on benchmarks. Pay attention to the various studies that have been released (peer reviewing OpenAI’s studies has been hard, turns out). You’re not getting linear growth with more data. You’re not getting exponential growth (which I suspect is what the investors are wanting).

You’re getting small incremental improvements simply from the adding more data. There are some things that the AI companies that are doing to improve performance for certain queries (this is human reinforcement, as well as some “safety” models and mechanisms) - but the idea that you just keep feeding a foundational model more data and it suddenly becomes much better is a logical fallacy and there’s not a lot of evidence for it.

Oh, oh gods, no. It can be very biased. It was trained on content curated from the internet.

I cannot do any better description than this Bloomberg Article - it's amazing, but it covers how image generators tend to racially code professions.

I'm skeptical that is even possible. Google Tried and wound up making "Racially diverse WWII German soldiers". AKA Nazis.

Right, so how did they train these things?

Permalink to “Right, so how did they train these things?”

The shortest answer is “a whole bunch of copyrighted content that a non-profit scraped from the internet”. The longer answer is “we don’t actually fully know because OpenAI will not disclose what’s in their datasets”.

One of the datasets, by the way, is Common Crawl - you can block its scraper if you desire. That dataset is available for anyone to download.

If you’re an artist that had a publicly accessible site, art on Deviantart, or really anywhere else one of the bots can scrape, your art has probably been used to train one of these models. Now, they didn’t train the models on “the entire internet”, Common Crawl’s dataset is around 90 TB compressed, and most of that is…. Well, garbage. You don’t want that going into a model. Either way, it’s a lot of data.

If you were a company who wanted to get billions of dollars in investment by hyping up your machine learning model, you might say “this is just how a human learns to do art! They look at art, and they use that as inspiration! Exactly the same.”

I don’t buy that. An algorithm isn’t learning, it’s taking pieces of its training set and reproducing it like a facsimile. It’s not making anything new.

I struggle with this a bit too. One of my favorite art series is Marcel Duchamp’s Readymades - because it makes you question “what is art, really?”. Does putting a urinal on its side make it art? For me, yes, because Duchamp making you question the art is the art. Is “Hey Midjourney give me Batman if he were a rodeo clown” art? Nah.

Thus, OpenAI is willing to go to court to make a fair use argument in order to continue to concentrate the research dollars in their pockets and they’re willing to spend the lobbying dollars to ask forgiveness rather than waiting to ask permission. There’s a decent chance they’ll succeed. They’ll have profited off of all of our labor, but are they contributing back in a meaningful way?

Let’s explore.

Part 2, or “how useful are these things actually”?

Permalink to “Part 2, or “how useful are these things actually”?”

Recycling truck
Paul Vasarhelyi @ Shutterstock #78378802

Let’s start with LLMs, which the AI companies claim to be a replacement for writing of all sorts or (in the case of Microsoft) the cusp of a brilliant Artificial General Intelligence which will solve climate change (yeahhhhh no).

Remember above how LLMs take statistically likely tokens and start spitting them out in an attempt to “complete” what you’ve put into the prompt? How are the AI companies suggesting we use this best?

Well, the top things I see being pushed boil down to:

  1. Replace your Developers with the AI that can do the grunt work for you
  2. Generate a bunch of text from some data, like a sales report or other thing you need “summarized”
  3. Replace search engines, because they all kind of suck now.
  4. Writing assistants of all kinds (or, if you’re an aspiring grifter, Book generation machine)
  5. Make API calls by giving the LLM the ability to execute code.
  6. Chatbots! Clippy has Risen again!

There are countless others that rely on the Illusion that LLMs can think, but we’re going to stay away from those. We’re talking about what I think is useful here.

The Elephant in the Software Community: Do you need developers?

Permalink to “The Elephant in the Software Community: Do you need developers?”

Okay, there are so many ways I can refute this claim it’s hard to pick the best one. First off, “prompt engineering” has emerged as a prime job, and it’s really just typing various things into the LLM to try and get the best results (again, manipulating the statistics engine into giving you output you want. Non-deterministic output). That is essentially a development job; you’re using natural language to try to get the machine to do what you want. Because it has a propensity to not do that, though, it’s not the same as a programming language where it does exactly what you tell it to, every time. Devs write bugs, to be sure, but what the code says is what you’re going to get out the other end. With a carefully crafted prompt you will probably get what you want out the other end, but not always (this is a feature, remember?)

The folks who are financially motivated to sell you ever increasingly complex engines are incentivized to tell you that you can cut costs and just let the LLM do the “boring stuff” leaving your most high-value workers free to do more important work.

And you know what, because these LLMs were trained on a bunch of structured code, yeah, you probably can get it to semi-reliably produce working code. It’s pretty decent at that, turns out. You can get it to “explain” some code to you and it’ll do an okay (but often subtly wrong) job. You can feed it some code, tell it to make modifications, or write tests, and it’ll do it.

Even if it’s wrong, we’ve built up a lot of tooling over the years to catch mistakes. Paired with a solid IDE, you can find errors in the LLMs code more readily than just reading it yourself. Neat!

I actually tried this recently when revamping the GW2 Assistant app. I’ll be doing a post on my experiment doing this soonish, but in the meantime let me summarize my thoughts (which are actually the second point):

An experienced developer knows when the LLM has produced unsustainable or dangerous code, and if they’re on guard for that and critically examine the output they probably will be more efficient than they were before.

Inexperienced developers will not be able to do that due to unfamiliarity and will likely just let the code go if it “works”. If it doesn’t work, they’re liable to get stuck for far longer than pair programming with a human.

Devin, the AI Agent that claims to be the first AI software engineer looks pretty impressive! Time for all software devs to take up pottery or something. I want you to keep an eye on those demos and what the human is typing into the prompt engine. One thing I noticed in the headline demo is that the engineer had to tell Devin 3 or 4 times (I kinda lost count) that it was using the wrong model and “be sure to use the right model”. There were also several occasions where he had to nudge it using specialized knowledge that the average person is simply not going to have. Really, go check it out.

Okay, so, we’re safe for a little bit right?

Well….no. I’m going to link to an article by Baldur Bjarnason: The one about the web developer job market. It’s pretty depressing, but it also summarizes my feelings well. Regardless of the merits of these AI systems (and I have a sneaking suspicion that the bubble’s going to pop sooner rather than later due to the intensity of the hype), CTOs and CEOs that are focused on cutting costs are going to reduce headcount as a money-saving measure, especially in industries that view software as a Cost Center. Hell, if Jensen Huang says we don’t need to train developers, we can be assured that the career is dead.

I think this is a long-term tactical mistake for a few reasons:

  1. I think a lot of the hype is smoke-and-mirrors, and there’s no guarantee that it’s going to be orders-of-magnitude better.
  2. We’ll make our developer talent pool much smaller, and have little to no environment for Juniors to learn and grow, aside from using AI assistants to do work.
  3. Once the cost of using AI tools increases, we’ll be scrambling to either rehire devs at deflated cost, or we’re going to try and wrangle less power hungry models into doing more development things.

Neat.

This LLM wish fulfillment strategy is essentially “I don’t have time to crunch this data myself, can I get the AI to do it for me and extract only the most important bits”. The shortest answer is “maybe to some degree of accuracy”. If you feed it a document, for example, and ask it to summarize - odds are decent that it’ll both give you a relatively accurate summary (because you’ve increased the odds that it’ll produce the tokens you want to see in said document) that will also contain degrees of factual errors. Sometimes there will be zero factual errors. Sometimes there will be many. Whether those are important or not depends entirely on the context.

Knowing the difference would require you to read the whole document and decide for yourself. But we’re here to save time and be more productive, remember? You’re not going to do that. You’re going to trust that the LLM has accurately summarized the data in the text you’re giving it.

BTW, by itself an LLM can’t do math. OpenAI is trying to overcome this limitation by allowing it to run Python code or connect to Wolfram Alpha but there are still some interesting quirks.

So, you trust that info, and you take it to a board presentation. You’re showcasing your summarized data and it clearly shows that your Star Wars Action Figure sales have skyrocketed. Problem is you’re an oil and gas company and you do not sell Star Wars action figures. Next thing you know, you’re looking like an idiot in front of the board of directors and they’re asking for your resignation. Or, worse, the Judge is asking you to produce the case law your LLM fabricated, and now you’re being disbarred. Neat!

Remember, the making shit up is a feature, not a bug.

But wait! We can technology our way out of this problem! We’ll have the LLM search its dataset to fact check itself! Dear reader, this is Retrieval Augmented Generation (RAG). Based on nothing but my own observations, the most common technique I’ve seen for this is doing a search first for the prompt, injecting those results into the context window, and then having it cite its sources. That can increase the accuracy by nudging the statistics in the right direction by giving it more text. Problem is, it doesn’t always work. Sometimes it’ll still cite fake resources. You can pile more and more stuff on top (like doing another check to see if the text from the source appears in the summary) in an ever-increasing race to keep the LLM honest but ultimately:

LLMs have no connection to “truth” or “fact” - all the text they generate are functionally equivalent based on statistics

RAG and Semantic Search are related concepts - you might use a semantic search engine (which attempts to search on what the user meant, not necessarily what they asked) to retrieve the documents you inject into the system.

The other technique we really need to talk about briefly is Reinforcement Learning from Human Feedback (RLHF). This is “we have the algorithm produce a thing, have a human rate it, and then use that human feedback to retrain / refine the model”.

Two major problems with this:

  1. It only works on the topics you decide to do it on, namely “stuff that winds up in the news and we pinky swear to ‘fix’ it”.
  2. It’s done by an army of underpaid contractors.

You’d be surprised just how much of our AI infrastructure is actually Mechanical Turks. Take Amazon Just-walk-out, for example.

What we wind up doing is just making the toolchain ever more complicated trying to get spicy autocomplete to stop making up “facts”, and it might have just not been worth the effort in the first place.

But that’s harder to do these days because:

Google’s been fighting a losing battle against “SEO Optimized” garbage sites for well over a decade at this point. Trying to get relevant search results amidst the detritus and paid search results has gotten harder over time. So, some companies have thought “hey! Generative AI can help with this - just ask the bot (see point #6) your question and it’ll give you the information directly”.

Cool, well, this has a couple of direct impacts, even if it works. Remember those hallucinations? They tend to sneak in places where they’re hard to notice, and its corpus of data is really skewed towards English language results. So, still potentially disconnected from reality (but usually augmented via RAG), but how would you know? It’s replaced your search engine - so are you going to now take the extra time to go to the primary source? Nah.

Buuuut, because Generative AI can generate even more of this SEO garbage at a record pace (usually in an effort to get ad revenue) we’re going to see more and more of the garbage web showing up in search. What happens if we’re using RAG on the general internet? Well, it’s an Ouroboros of garbage, or, as some folks theorize, Model Collapse.

The other issue is that if people just take the results the chat bot gives them and do not visit those primary sources, ad revenue and traffic to the primary sources will go down. This disincentivizes those sources from writing more content. The Generative AI needs content to live. Maybe it’ll starve itself. I dunno.

But it’ll help me elevate my writing and be a really good author right?

Permalink to “But it’ll help me elevate my writing and be a really good author right?”

I’ve been too cynical this whole time. I’m going to give this one a “maybe”. If you’re using it to augment your own writing, having it rephrase certain passages, or call out to you where there are grammar mistakes, or any of that “kind” of idea more power to you.

I don’t do any of that for two reasons, one is practical, the other highlights where I think there’s an ethical line:

  1. I’m not comfortable having a computer wholesale rewrite what I’ve done. I’d rather be shown places that can improve, see some other examples, and then rewrite it myself.
  2. There’s a pretty good chance that the content it regurgitates is copyrighted, and we’re still years out from knowing the legal precedent.

The AI industry has come up with a nice word for “model regurgitates the training data verbatim”. Where we might call it “plagiarism” they call it “overfitting”.

Look, I don’t want to be a moral purist here, but my preferred workflow is to write the thing, do an editing pass myself, and then toss the whole thing into a grammar checker because my stupid brain freaking loves commas. Like, really, really, loves them. Comma.

I do this with a particular tool: Pro Writing Aid. It’s got a bunch of nice reports which will do things like “highlight every phrase I’ve repeated in this piece” so that I can see them and then decide what to do with them. Same deal with the grammar. I ignore its suggestions frequently because if I don’t, the piece will lose my “voice” - and you’ll be able to tell.

They, like everyone else, have started injecting Gen AI stuff into their product, but for me it’s been absolutely useless. The rephrase feature hits the same bad points I mentioned earlier. They’ve also got a “critique” function which always issues the same tired platitudes (gotta try it to understand it, folks).

This raises another interesting point about the people investing in Generative AI heavily. One of those companies is Microsoft. A company who makes a word processor. The parent of clippy themselves. They could have integrated better grammar tools into their product. They could have invested more in “please show me all the places where I repeated the word ‘bagel’”. They didn’t do this.

That makes me think that they didn’t see the business case in “writing assistants”, and why Clippy died a slow death.

Suddenly, though, they have a thing that can approximate human writing and suddenly there’s a case and a demand for “let this thing help you write”. I feel like they’re grasping at use cases here. We stumbled upon this thing, it’s definitely the “future”, but we don’t…quite….know….how.

I want to take a second here to talk about a lot of what I’m seeing in the business world’s potential use cases. “Use this to summarize meetings!” or “Use this to write a long email from short content” or “Here, help make a presentation”.

After all one third of meetings are pointless, and could be an email! I want to also contend that many emails are pointless.

Essentially what you’re seeing is a “hey, busywork sucks, let’s automate the busywork”. Instead of doing that, why not just…not do the busywork? If you can’t be bothered to write the thing, does it actually have any value?

I’m not talking about documentation, which is often very important (and should be curated rather than generated), but all those little things that you didn’t really need to say.

If you’re going to type a bulleted list into an LLM to generate an email, and the person on the other end is going to just use an LLM to summarize, lossily, I might add, why didn’t you just send the bulleted list?

You’re making more work for yourself. Just… don’t do that?

Let’s give it the ability to make API Calls

Permalink to “Let’s give it the ability to make API Calls”

Right, so one of the fun things OpenAI has done for some of their GPT-4 products is to give it the ability to make function calls, so that you can have it do things like:

  • Book a flight
  • Ask what the next Guild Wars 2 World Boss is
  • Call your coffee maker and make it start
  • Get the latest news
  • Tie your shoes (not really)

And so on. Anything you can make a function call out to, you can have the LLM do!

It does this by being fed a function signature, so it “knows” how to structure the function call, and then runs it through an interpreter to actually make the call (cause that seems safe).

Here’s the…minor problem. It can still hallucinate when it makes that API call. So, say you have a function that looks like this: buyMeAFlight(destination, maxBudget) and you say to the chatbot “Hey, buy me a flight to Rio under $200”. What the LLM might do is this: buyMeAFlight("Rio de Janeiro", 20000). Congrats, unless you have it confirm what you’re doing you just bought a flight that’s well over your budget.

Now, like all other Generative AI things, there are techniques you can use to increase the accuracy. Making just the perfect prompt, having it repeat output back to you, asking “are you sure”, telling it that it’s a character on star trek. You know, normal stuff.

Alternatively you could just... use something deterministic, like, I don’t know, a web form or any of the existing chat agent software we already had.

Sidebar: Apparently OpenAI has introduced a “deterministic” mode in beta, where you provide a seed to the conversation to get it to reliably reproduce the same text every time. Are you convinced this is a random number generator yet?

And on that note, let’s talk about our obsession with chatbots.

So the “killer” application we’ve come up with, over and over, is “let’s type our question in natural language and it does a thing.” I honestly don’t understand this on a personal level - because I don’t really like talking to chatbots. I don’t want to say “Please book me a flight on Friday to New York” and then forget about it. I want to have control over when I’m going to fly.

Do large swaths of people want executive assistants to do important things like cross-country travel?

Not coincidentally, I really struggle with that kind of delegation and have never really made use of an executive assistant personally.

We’ve decided that the best interface for doing work is “ask the chatbot to do things for you” in the agent format. This is exactly the premise of the Rabbit R1 and the Humane Ai Pin. Why use your phone when you can shout into a thing strapped to you and it’ll do…whatever you ask. Perhaps it’ll shout trivia answers at you.

But guess what, my phone can already do that. Siri’s existed for years and like, I hardly use it. It’s not because it’s not useful. It’s because I can do what I want without shouting at it. In public. For some reason.

We do need to talk about accessibility. One of the things that AI agents would be legitimately useful for is for those folks who cannot access interfaces normally whether that’s situationally (driving a car), or temporarily / permanently (blind, disabled).

If we can use LLMs to get better accessibility tech that is reliable, I’m all for it. Problem is that the companies pushing the technology have a mixed track record on doing accessibility work, and I’m concerned that we’ve decided that LLMs being able to generate text means we can abdicate responsibility for doing actual accessibility work.

Like many other things in the space, we’ve decided that “AI” is magic, and will make things accessible without having to do the work. I mean, no. That’s not how it works.

Remember back to the beginning of this article where I talked about other Machine Learning Models? I think that’s the space where we’re going to make more accessibility advances, like the Atom Limb which uses a non-generative model to interpret individual muscle signals.

Still with me?

If I had to summarize my thoughts on all of the above it’s that we’ve stumbled upon something really cool - we’ve got an algorithm that can create convincing looking text.

The companies that have the resources to push this tech seem to be scrambling for the killer use-case. Many companies are clamoring for things that let them reduce labor costs. Those two things are going to result in bad outcomes for everyone.

I don’t think there’s a silver bullet use case here. There are better tools already for every use case I’ve seen put forward (with some minor exceptions), but we’re shoving LLMs into everything because that’s where the money is. We’re chasing a super-intelligent god that can “solve the climate crisis for us” by making the climate crisis worse in the meantime.

If you were holding NVDA stock, something something TO THE MOON. They’ve been making bank off of every bubble that needs GPUs to function.

This feels exactly like the Blockchain and Web3 bubbles. Lots of hype, not a lot of substance. We’re tying ourselves in knots to get it to not “hallucinate”, but like I’ve repeated over and over again in this piece the bullshit is a feature, not a bug. I recommend reading this piece by Cory Doctorow: What Kind of Bubble is AI? It’ll give you warm fuzzies. But it won’t.

Midjourney, Sora, all those things that can fake voices and make music. We’ve got a big category of things that are, charitably “art generators”, but more realistically “plagiarism engines”.

This section is going to be a lot shorter. Let me summarize my feelings:

  • If you’re using one of these things for personal reasons, making character art for your home D&D game, or other things that you’re not trying to profit from - go for it. I don’t care. I’d rather you not give these companies money but I don’t have moral authority here.
    • I’ve used it for this too! I’m not exempt from this statement.
  • If you’re using AI “art” in a commercial product, you don’t have an ethical defense here (but we’ll talk about business risk in a sec). The majority of these models were trained on copyrighted content without consent and the humans who put the work in are not compensated for it.

I personally don’t find all of the existing AI creations that inspiring, other than how neat it is we’ve gotten a neural network to approximate images in its training set. Some of the things it spits out are “cool” and “workable” but I just don’t like it.

Hey, I do empathize with the diffusion models a bit though. Hands are hard.

As I mentioned earlier in the post, as far as we can tell, the art diffusion models were trained on publicly viewable, but still copyrighted content.

If for some reason you’re a business and you’re reading this post: That’s a lot of business risk you’d be shouldering. There are multiple different lawsuits happening right now, many of them on different lines, and we don’t actually know how that’s going to go. Relatedly, AI Art is not copyrightable, so that’s…probably a problem for your business especially if you’re making a book or other art-heavy product. At best you can do is treat it like stock art, where you don’t own the exclusive rights to the art, and you’re hoping you don’t get slapped with liability in the future.

So, if you’re using an AI Art model in your commercial work, these are all things you have to worry about.

This is where, and I cannot believe I am saying this, I think Adobe is playing it smart. They’ve trained Firefly on Art they’ve licensed from their Adobe Stock art platform and are (marginally) compensating artists for the privilege. They have also gone so far as to offer to guarantee legal assistance to enterprise customers. If you’re a risk averse business, that’s a pretty sweet deal (and less ethically concerning - though the artists are getting pennies).

The rest of them? You’re carrying that risk on your business.

But what if your business happens to be “crime”?

AI companies seem hell bent on both automating the act of creation (devaluing artistry and creativity in the process) and also making it startlingly easy for fraudsters to do their thing.

Lemme just… link some articles.

So, you create things that can A) Clone a person’s voice, B) Imitate their Likeness, and C) Make them say whatever you want.

WHAT THE HELL DID YOU THINK WAS GOING TO HAPPEN? WHAT PUBLIC GOOD DOES THAT SERVE?

I dunno about y’all, but I’m okay not practicing Digital Necromancy (just regular, artisanal necromancy).

The commercial business for these categories of generative AI are flat out fraud engines. OF COURSE criminals are going to use this to defraud people and influence elections. You’ve made their lives so much easier.

Hey, I guess we can take solace in the fact that the fraud mills can do this with fewer employees now. Neat.

But Netflix Canceled My Favorite Show and I want to Revive it

Permalink to “But Netflix Canceled My Favorite Show and I want to Revive it”

This is also known as the “democratizing art” argument. First thing I would like to point out is that art is already democratized? It’s a skill. That you can learn. All you need to do is put in the time. It’s not a mystical talent that only a select few possess.

Artists are not gurus who live in the woods and produce art from nothing, and the rest of us are mere drones who are incapable of making art.

So in this case “democratization” really means “can make things without putting in the effort”. The result of that winds up being about as tepid as you might imagine.

Now, there’s a question in there of if a person is having to work all the time to simply live, won’t this enable them to “make art”? There’s another way to fix that, by reducing the amount of work they need to do to simply exist, but nah, we’re gonna automate the fun parts.

Hey, awesome artists who are making things with AI tools but are using it as a process augmenter - all good. I’m not talking to you. I’m talking to the Willy Wonka Fraud Experience “entrepreneurs”

But you know what, I’m not even really that concerned with people who want to make stuff on their own that is for their own enjoyment. I don’t think the result is going to be very good, and I’d rather have more people creating good stuff than fewer, but hey more power to ya.

Another aside: I really do not want to verbally talk to NPCs in games. I play single-player games to not talk to people. I don’t want to be subjected to that in the name of “more realistic background dialog”.

It’s just not going to work out like you think it will. What’ll actually happen is:

The primary thing I suspect is going to happen with the AI “art” is back to the cost-cutting efforts. Where you might have used stock art before, or a junior artist, you’re going to replace that with Dall-E.

For marketing efforts, that’s not an immediate impact. Marketing content is designed to be churned out quickly, and shotgunned into people’s feeds in an effort to get you to buy something or to feel some way, etc. I don’t think those campaigns are going to be as effective as the best campaigns ever, but eh, we’ll see I guess.

The most concerning uses are going to be the media companies that are going to replace assets in video games and movies. Fewer employees, lower budgets, and … dare I say … lower quality.

You see, diffusion models don’t let you tweak them, yet (although, who knows, maybe if we start doing deterministic “seeds” again we’ll get somewhere with how Sora functions). They also have a propensity to not give you what you asked for, so, yeah, let’s spend billions of dollars trying to fix that.

So, at risk of trying to predict the future (which I’m definitely bad at), I think we’re going to gut a swath of creatives, devalue their work, and then realize that “oh, no one wants this generative crap”. We’ll rehire the artists at lower rates and we’ll consolidate capital into the hands of a few people.

Meanwhile, we’ve eliminated the positions where people would traditionally learn skills, so you won’t be able to have a career.

Because…

We live in a society that requires money to function. Most of us sell our labor for the money we need to live.

The goal of companies is to replace as many of their workers as they can with these flawed AI tools. Especially in the environment we find ourselves in where VC money needs to make a return on investment now that the Zero Interest Rate Phenomenon (ZIRP) is finished.

Now, “every time technology has displaced jobs we’ve made more jobs” is the common adage. And, generally, that’s true. However, the main fear here isn’t that we won’t be working, it’s that it’ll have a deflationary impact on wages and increase income inequality. CEO compensation compared to average salary has increased by 1460% since the 70’s, after all.

What I think is different about previous technological advances (but hey, the Luddites were facing similar social problems) is that we’re in a situation where the amount of capital is being invested in the hands of a few companies, and only a very few of them have the resources to control this new technology. I don’t think they’re altruists.

This is not a post-scarcity Star Trek future we’re living in. I wish we were. I’m sorry.

Right. Uh. I’m not sure I can, but here are some things I’d like to see:

  • There are a number of really small models that can run on commodity hardware, and with enough tuning you can get them to give you comparable results to what you’re getting on some of the larger models. Those don’t require an ocean of water to train or use, and run locally.
  • We’re going to see more AI chips, that’s inevitable, but the non-generative models are going to benefit from that too. There’s a lot of interesting work happening out there.
    • I’m also pretty cool with DLSS and RSR for upscaling video game graphics for lower-powered hardware. That’s great.

I honestly hope I’m wrong and that the fantastical claims about AI solving climate change are real… but the odds of that are really bad.

This is the longest post I think I’ve ever written, we’re well over 7,000 words. I have so many thoughts it’s hard to make them coherent.

Perhaps unsurprisingly, I’ve not used Generative AI (or AI of any kind) for this post. I’ve barely even edited it. You’re getting the entire stream of consciousness word vomit from me.

Call me a doomer if you want, but hey, if you do, I’ve got a Blockchain to sell you.

Generative AI for Coding

AI Clacky Keyboard Tests
Owlie Productions @ Shutterstock #2372381633

Update 7-13-24: The tl;dr of this post is "You shouldn't use these, it's not worth it", in case that's unclear.

What happens when you get an extreme “GenAI Skeptic” and shove him in front of an LLM coding assistant? This, turns out.

This is a follow up on my last post, which was an epic rant about Generative AI. In that, I mentioned that while I’m generally skeptical about GenAI replacing developers (for a few different reasons, but also because they can’t do what their promoters say they can), I do think it’s one of the use cases where it could actually be useful as a productivity augment for the “write code” portion of the project (if it worked better and wasn't setting the planet on fire to do it).

I don’t think it’s going to actually be able to replace developers. I think that the AI companies want you to think it’s going to be able to replace developers, but they’re not currently capable of doing so. Watch this video on Debunking Devin by Internet of Bugs, it’s a good watch.

Likewise, the “writing code” portion of software development isn’t the main point of software development, and the LLM can’t do all the fiddly human bits of creating a software product.

Before I get into how I’ve tested this myself I want to call out a couple of posts from people I respect that cover GenAI more thoughtfully and less “cathartic rant-centric” than I was.

First up, Molly White wrote AI isn’t useless. But is it worth it?. Go read it, it’s great. In it, she mentions something I’m glad she did: there are other, better, tools that use way less energy than existing tools for proofreading / editing / grammar checking. But I basically agree with the entire post (and I think it makes a lot of the same points I did, more elegantly and thoughtfully).

The other one is a Guardian Article (which….wow I’m linking to the Guardian), that talks about the AI bubble we’re definitely in.

The tl;dr or “I’m going to have GPT summarize this for me"

Permalink to “The tl;dr or “I’m going to have GPT summarize this for me"”

I want to emphasize that these are my opinions, and if you find LLMs for coding personally useful, that’s okay. I don't think you should use them, though, and I'm getting to the point where if you mention "I used GPT for..." it's more likely that I'm going to not give the rest of your argument much weight.
I also want to call out that I do not use Generative AI for my writing. There are other tools for editing and grammar checking and thesauruses and so on.

For me, there’s some utility in how these things operate. That utility is variable and hard for me to properly quantify. Sometimes, it’s a time save. Sometimes, it’s a time sink. If I had to guess, I’d say it’s a net time-save right now, but that time save is not nearly enough to offset the environmental and social costs.

“But Alex, these will get exponentially better, and will eventually do everything for you”, you might be saying. I’m not trying to set up a straw man. This is what the AI companies are selling. I don’t actually think this is true. My prediction is that we’re already reaching the end of the exponential growth curve, and the amount of utility we can get out of LLMs will plateau.

And look, even if I’m wrong, the change of pace here is so fast - you’re not going to be missing out if you don’t adopt LLMs for coding right now. If you’re a business owner, I’d argue that waiting a bit longer makes more business sense because I suspect once the free money runs out and these folks need to turn a profit the cost is going to go way up, and you’re going to have to recalculate your costs again.

So maybe it’ll be more useful (for me) eventually. Maybe they’ll solve the energy requirements and you can run one of these models locally.

I’m just one guy talking about his own experiences with a Coding LLM. I obviously think I’m right, otherwise I would be singing a different tune, but I want to drop a quote here:

IOW, everything written about LLMs from the perspective of a single practitioner can be dismissed out of hand. The nature of LLMs makes it impossible to distinguish signal from noise in your own practice. - Baldur Bjarnason, The Intelligence Illusion

That is to say, any individual account, whether positive or negative for LLMs, is inherently biased. (See also - You should not be using LLMs)

If you want a counterpoint, go have a look at Simon Willison’s blog, which Molly linked to in her article. I disagree with his assessment of the ethics and the inevitability, but go have a look and learn for yourself.

Sidebar: Even his posts that have nothing to do with Generative AI have started to include statements like "So I asked GPT-4o to help me...".

Okay, that out of the way, let’s talk about the Coding Experiment.

I set a couple of rules for myself to try and make this experiment have some constraints so that I can emulate how I’d expect a competent coding assistant to function.

  1. Minimal “prompt engineering”. The tool can inject whatever context it wants, but I’m going to use the tool as the marketing says I should be able to.
    1. Likewise, if I need to type out more words to describe what I want than it’d have taken to simply write the code…. that’s not great.
  2. Working on a well-represented language: Typescript is well-represented in GPT’s dataset.
  3. Real project I’m working on.
  4. No googling, only doing what the Assistants tell me to do.

Here are the tasks I’m going to accomplish:

  1. Replace the Deprecated ‘request’ module with Axios in the http class.
  2. Fix the failing unit tests and refactor them to use async / await
  3. Fix a problem with getDeeperInfo related to closures / scope.

I’ve chosen 3 different projects to run this test on:

  1. Zed with GPT-4-turbo
  2. Continue.dev (Extension) on VS Code with GPT-4-Turbo
  3. Github Copilot on VS Code

The first two are because I want to do a “control” for the LLM (though they do use slightly different versions). The last one is just for a commercial-off-the-shelf “optimal” experience.

I don’t know what extra context the tools are injecting before sending data to their LLM counterpart, but hopefully by doing two with the same model we’ll get a decent comparison between the two.

I’m not really looking to rate any of these things as a clear “winner”, but I did want to see if any of these clearly outperformed each other.

Copilot had the weirdest behavior of them all, it would frequently not update the code inline even when I told it to. It also was the only one that flat out started removing braces, leading to a fun little compile error.

But, overall eventually each of the three were able to assist with the tasks. They all performed “best” at transforming existing code - all three of them were able to turn promises into async/await without too much trouble.

All three of them had some issues with creating more code than was necessary, or generating code that doesn’t work. They all did a decent job of summarizing code (with some fun little inaccuracies), and usually were able to help me spot things like a missing return statement that an IDE could notice, but is frequently not configured to notice.

Like I said in the tl;dr - these things were fiddly and inconsistent. Frequently, the built-in IDE features we’ve already had were much better. What you're being sold right now is a future where this works better, which would be "fine" if the hype cycle weren't also selling these tools as a "Developer replacer" to execs (shout out to Copilot for slapping warnings of "this doesn't replace human effort" all over the actual tool... but it's not going to be enough).

So the first test actually took the longest, largely because this was the test where I had no prior experience with the issues that I need to fix. The second and third tests got progressively faster as I knew where to look for things, but the LLM got a little… “spicy”.

If you want to have a look at the test, here’s the video:

First off, it did a fine job replacing requests with axios, though it (and my undercaffeinated brain) had some issues with the typings. When it came to actually call the API though, it introduced some mistakes with how the API was called. It’s arguable here that I’d have gotten a better result with a better prompt, but looking back to the rules, I shouldn’t have to type a paragraph.

On the second test, it did fine, though I did forget to have it do an async/await in the video. I did try it later, and it worked fine - basically the same as the others. This was probably the most consistent, but it’s also one of the things where I can complete the task / update in about 30-45 seconds and the LLM does it in around 20 seconds (inclusive of typing the prompt).

Sidebar: my best time using liberal copy/paste in that was around 16 seconds. This is what I’m taking about with variable time save.

For the Third test - I wound up not needing the LLM to help with it. It was a missing param on a couple of calls. So, it wound up being unnecessary.

Overall: This was “fine” - and it mostly stayed out of my way while I was doing the experiment, which I appreciate.

Right, so this one was the first of the two that use VS Code Plugins. Continue.dev is an open-source way to call all sorts of LLMs, including locally hosted LLMs. I’ve also given that a shot but the experience untethered from a GPU is not great.

For this test I thought I’d disabled the code autosuggestion, but for some reason the config didn’t stick and it was using…some…. llm’s free trial API for it? I honestly have no idea what happened there and it’s present in the video.

Speaking of video, the commentated test is here:

Like the others, this one did a pretty okay job at each of the tasks, but it had some standouts:

  1. The inline updates worked consistently and made an actually diff in the IDE, so I could accept or reject things individually. That made it a lot easier than Zed to see if the LLM had inserted something that I didn’t want (which happened a lot).
  2. Subjectively I think there were more hallucinations, but it’s the same LLM so that’s probably random.

It’s worth calling out here that all of these calls to the LLM are non-deterministic. If you were to set the temperature to 0 you could get more deterministic behavior but that would also destroy its “creativity” so no one really does that.

Like before, it completed task 1 in a very similar way to Zed, but I had to futz with the output more. It created a similar error when refactoring the actual callAPI method as well - which took a little bit longer to fix.

For task 2, it did just fine, and it was likewise able to eventually figure out there was a missing return statement.

Right, so Copilot was frustrating in more way than one.

Sidebar: Microsoft, I need you to get your shit together and stop naming different products the same thing please.

First off, the inline editing was extremely inconsistent. If you check out the video, you’ll encounter this immediately:

For the first several attempts, it just doesn’t work. It won’t make changes, or even accurately suggest what to do - it just craters and suggests I use the chat sidebar. Eventually it starts working and I assume that it’s looking for cues in the response from the LLM to do its inline replacements but it messes up more than once.

The chat is “fine”, in that it’s a bit verbose (how much money are we burning on extra tokens?) and it’s moderately helpful at times.

For problem 1, I wound up having to use the autosuggest to make the changes, after several failed attempts at the inline changes.

For problem 2, it worked fine - like the others it was able to both suggest the missing return (though it also added a bunch of unnecessary code) and to refactor the tests to be async.

So, for Github Copilot, I’d never used it before so I got to get a 30 day free trial! Hooray!

For Zed and Continue, I was using my OpenAI API key (which yes, I do have an OpenAI API key - I’ve spent about $10 on it total so far for experimentation). Have a look:

I spent $1.55 for approximately 1.5 hours of “light” usage (maybe we be generous and call it 2). If we assume that I’m using that level of usage across all “tasks” that the AI folks envision (like, I’m using Copilot for code and the other Copilot for Emails, because they’re shoving them into absolutely everything) - we can “extrapolate” 160 hours at roughly $0.75 per hour puts us at $120 of usage per month. Were I using GPT-3.5 you can pretty well divide that by 20 (ish) for a usage of $6/mo.

You might notice there’s a math problem here. Assuming Copilot is using GPT-4, and they charge $10/mo - they’re likely loss-leading a lot here. If they’re using a more efficient model….they’re probably still losing money on every subscription unless devs aren’t using features.

I suspect this is also the case for the Office 365 + Edge Browser + everything else - burning investment dollars to get everyone hooked.

I only mention this because I think they’re going to raise prices eventually, unless there’s some egregious advance in chip technology (which Nvidia is chasing), and when they do….

So yeah, to summarize what I said at the beginning, I’m just not getting enough value for these things to justify their use and the ethics and energy use bother me.

If we can get to a point where we’re not burning egregious amounts of energy and these models aren’t consolidating even more capital in the big tech companies, that might change. Maybe we'll get really advanced AI chips that let you run all of these locally on commodity hardware. I just think the current trend of "more power and bigger" is unsustainable.

I still also have some existential concerns with newer developers learning through the use of LLMs, when they'd be better served by really good curated documentation, but I'll leave that for another post.

11ty Conference Thoughts

The International Symposium on Making Web Sites Real good is now over and I watched it live. All of the talks are watchable up on Youtube.

Overall I really enjoyed all of the talks, especially the ones that made me think a lot more about what the internet is about. The whole conference is well worth watching, but I'm just going to call out a few big highlights that I especially enjoyed.

Here's the full video:

And the highlights:

This was an amazing talk all around, and discusses how the web was originally designed to protect content over other concerns and how styling came to exist in a way that respects that end goal. The overall discussion around the history of the web and how browsers preserve and protect the content and then how that interacts with authorial intent on a webpage is just fascinating. There are a bunch of links to old web RFCs as well.

DIGITAL FRONTIERS, INDIEWEB COWBOYS, AND A PLACE ONLINE TO CALL YOUR OWN

Permalink to “DIGITAL FRONTIERS, INDIEWEB COWBOYS, AND A PLACE ONLINE TO CALL YOUR OWN”

Henry was an extremely enthusiastic and energetic presenter and clearly loves the stuff he's talking about. He goes into a speedrun of IndieWeb concepts (which is great because he also does Elden Ring challenge runs), and I really recommend jotting down some of the things he talks about in here. Here are some links to some of the concepts / tech he talks about in the talk.

I especially loved how all of these concepts work together and don't require a major commercial entity to work. You can roll everything yourself. I especially like IndieAuth (even though I do prefer other authentication techniques).

This talk was really fun to watch. Dan talks at length about creating a meta-web interactive story which spans 40 websites, at least one Instagram page, and a bunch of different plot lines of a fictional city. It gives Welcome to Nightvale vibes, and is really reminiscent of some of the best parts of the early web.

That's a lot of domains to keep registered!

It's also thought provoking for me because they did make use of Generative AI (apparently more than they thought they would have) in order to make the site happen. Dan didn't go into detail in the talk about just how they did that, but it would appear that it was largely in the image generation space with some of the "general site copy" being LLM generated.

On one hand, I appreciate that making a project of this size would have been impossible for a team of 2 writers to do previously if you don't want to resort to stock art or other licensed images. Being able to generate fictional people enhanced the story they wanted to tell. They managed to create art, where the focus of the art is the story and not the imagery. The imagery does help sell the fiction of the story, especially when it's a little "weird".

On the other hand, I kinda still think this is a borderline unethical use of the technology since I think it's using a model trained on copyrighted data. This is mitigated somewhat from this not being a directly commercial endeavor, but it still kinda makes me feel weird, ya know?

Note, I'm making an assumption here, I can't find what models they used to do this but I think it's Midjourney rather than something like Firefly.

I'd love to see a discussion of the ethical considerations of undertaking a project like Question Mark, Ohio using generative AI.

COME TO THE LIGHT SIDE: HTML WEB COMPONENTS

Permalink to “COME TO THE LIGHT SIDE: HTML WEB COMPONENTS”

Chris talks about a concept called "HTML Web Components" which is essentially "use regular HTML until it hits its limits and then enhance those with web components.

If you know anything about me, you know that I love web components. Love them. I love the concepts behind this talk, and it really gives the vibes of "we used to do this with jQuery but we have better tools now" vibes. Definitely worth a watch.

Chris is an excellent presenter and sticks to VanillaJS rather than using something like Stencil, which is great for learning the actual component API. Also the cadence and tenor of his talk is quite soothing.

This is just fascinating - he dives into the history of Chinese Type Systems and ties them to the web beautifully. I study Japanese (which uses a lot of Chinese characters) and so it was fun to understand how these fonts get rendered, how you optimize for them, and a bunch of other neat things I'd never have thought of.

Go give it a watch.

That's it! I enjoyed the whole conference and really appreciate the community / 11ty / CloudCannon for putting it together. Thanks to all the presenters for their time, and thanks to the chat for awesome commentary.

Microsoft’s Copilot+ Recall is a Horrible Idea

Shutterstock evil cube
80's Child @ Shutterstock #1235192320

…and you should disable it. Or not buy a Copilot+ PC.

I’m going to cite every source I’ve got on this, the story out of Microsoft is changing and being “clarified” as this goes on, so I’ll do my best to keep this updated and as accurate as possible.

Update - October 15th, 2024: It's baaaack. Microsoft details security/privacy overhaul for Windows Recall ahead of relaunch. There's also some chatter about it being enabled for every Windows 11 PC, not just Copilot+ ones. Check out this video:

Update - June 13th, 2024: It just keeps getting better and better. Ars Technica says Microsoft is in full damage control mode with just 2 days to go until rollout.

Update - June 7th, 2024: The Verge is now reporting that Microsoft is going to make some changes after the uproar. I'm not sure this matters, I don't think Microsoft has anyone's trust that they won't make changes in the future to get at all that juicy user-generated data. But good on them for doing something.

Update - June 2nd, 2024: Oh good, Nvidia is partnering to enable this on even more computers. It's going to be basically everywhere soon whether you want it or not. Nice. Apropos of nothing I've moved my Framework to running Bazzite. It's working great so far. Haven't tried an eGPU yet.

Update - May 31st, 2024 Kevin Beaumont has posted a lengthy Q/A style post about this which is very good on his blog: Stealing everything you’ve ever typed or viewed on your own Windows PC is now possible with two lines of code — inside the Copilot+ Recall disaster. You should go give that a read. The one thing I call out is I think the wording implies it's also acting as a keylogger, but I don't think it is necessarily, but it is screenshotting often enough that whatever you type is going to wind up in a screenshot.

Okay, here it is again. Another post about Generative AI. Or, rather, "bad ideas brought about by the generative AI hype cycle". Really, I would rather not be spending more time writing about all of this, but it just keeps finding me somehow. I’m so tired.

Anyway, let's talk about Copilot+ Recall (apparently it’s just “Recall” but that’s really difficult to web search, so I’m going to use Copilot+ in front of it). The idea behind it is that everything you do on your computer, will be screenshotted every {n} seconds and stored by the Neural Processing Unit (NPU) somewhere on your computer. It'll do Optical Character Recognition (OCR) on these screenshots and do some other magical "AI goodness" to enable you to later query for anything you did within the past {n} months (configurable, based on how much storage you want to use). Here's the official Microsoft documentation on the feature.

Sidebar, apparently this is enableable on non-Copilot+ computers, but you have to go out of your way to do it.

You know what else records and stores everything you do on your computer? A rootkit. That's right, this thing that Microsoft is installing and apparently enabling by default on new Copilot+ machines, is behaving exactly like a computer virus wants to.

During setup of your new Copilot+ PC, and for each new user, you're informed about Recall and given the option to manage your Recall and snapshots preferences. If selected, Recall settings will open where you can stop saving snapshots, add filters, or further customize your experience before continuing to use Windows 11. If you continue with the default selections, saving snapshots will be turned on (emphasis mine). — Microsoft

How helpful!

Microsoft's main "defense" has thus far been "but it only stays on your computer and never leaves the network". This totally ignores the fact that viruses that want to gain access to that process could just do that itself, regardless of Microsoft's wishes.

This thing will be an incandescent target for hackers. You'd think that Microsoft would know this, but it seems like in their rush to "win" the AI hype race they've cut some corners (which is unsurprising given their recent track record on security).

Recall's security is entirely based on “but it stays local”

Permalink to “Recall's security is entirely based on “but it stays local””

But “stays local” is not the same as “secure”. The only fool-proof security is to not store the thing in the first place, but here we are.

Let’s look at some of the corners they’ve cut.

First off, according to Kevin Beaumont, the NPU takes the text it extracts from the images and stores it in a user-readable sqlite database. This is very convenient for searching! This is also very convenient for any malicious process that happens to be running as you to ship off.

Guess what else it does (or rather doesn't do): obfuscate passwords or other sensitive information:

"Note that Recall does not perform content moderation. It will not hide information such as passwords or financial account numbers. That data may be in snapshots that are stored on your device, especially when sites do not follow standard internet protocols like cloaking password entry." — Techradar

Riiiight, so I'm guessing it's not going to obfuscate things like, I dunno, your bank's website? Possibly showing your full account details?

Are there… any other things you do on your computer that you'd rather not be stored for however long Recall wants to store the screenshots? Nothing? Anyhow, maybe you can take some small solace in the fact that apparently Microsoft does know how to process content and won't store DRM'd things:

"Recall also does not take snapshots of certain kinds of content, including InPrivate web browsing sessions in Microsoft Edge. It treats material protected with digital rights management (DRM) similarly; like other Windows apps such as the Snipping Tool, Recall will not store DRM content." — From that same Techradar article

… Cool. Now, Microsoft claims the following about other browsers:

Recall won’t save any content from your private browsing activity when you’re using Microsoft Edge, Firefox, Opera, Google Chrome, or other Chromium-based browsers.

Which was not the case when they first announced the idea, near as I can tell. So I guess they walked back from the “this only works in Edge”.

Microsoft Promises you can Manually exclude apps

Permalink to “Microsoft Promises you can Manually exclude apps”

Users can pause, stop, or delete captured content and can exclude specific apps or websites. — Ars Technica Article

But you’ve put the burden on the user to A) Know this is happening and B) Actually manage to catch everything they don’t want included.

Microsoft says, “Trust us, we’ve got Secure Core and Pluton processors!”

Permalink to “Microsoft says, “Trust us, we’ve got Secure Core and Pluton processors!””

The security protecting your Recall content is the same for any content you have on your device. Microsoft provides many built-in security features from the chip to the cloud to protect Recall content alongside other files and apps on your Windows device.
Secured-core PC: all Copilot+ PCs will be Secured-core PCs. This feature is the highest security standard for Windows 11 devices to be included on consumer PCs. For more information, see Secured-core PCs.
Microsoft Pluton security processor will be included by default on Copilot+ PCs. For more information, see Microsoft Pluton. — Microsoft

Microsoft, buddy. None of those things matter if you trick the user into running something in user space because you’ve granted the user access to that database.

None of those things matter if someone with access to the computer wants to go looking back through your history forever.

By the way, if you go look at the Pluon processor, it doesn’t mention a word about Windows Home edition, so I dunno if every Copilot+ machine is going to come with a Windows Pro license or what?

But surely these things are encrypted on the device using BitLocker? Right?

Permalink to “But surely these things are encrypted on the device using BitLocker? Right?”

And on that previous note, only if you have a Business or Pro license?

This one is a little tricky because Microsoft is apparently being a little vague, but by all accounts it looks like Home users get to have unencrypted screenshots just sitting on their laptop. Fun!

Sidebar: To see those screenshots, the user needs to be able to decrypt them so…. Again… virus. Running as the user. Can see them.

Maybe all Copilot+ machines come with Windows Pro? I just dunno.

Who else might have access to your computer?

Permalink to “Who else might have access to your computer?”

Many other people who are better able to speak to this problem have pointed out that this is rife for abuse by abusive partners.

In fact, Recall seems to only work best in a one-device-per-person world. Though Microsoft explained that its Copilot+ PCs will only record Recall snapshots to specific device accounts, plenty of people share devices and accounts. For the domestic abuse survivor who is forced to share an account with their abuser, for the victim of theft who—like many people—used a weak device passcode that can easily be cracked, and for the teenager who questions their identity on the family computer, Recall could be more of a burden than a benefit. — Malwarebytes

So, yeah. That’s… great.

Some of them definitely will. Look, your corporate computer is already watching what you do for good reason. A corporation needs to know if their systems are being used for nefarious, illegal, or other things that will be a problem for them (exfiltrating product secrets, for example).

This goes to a whole other level. On one hand, this is the ultimate forensic “we need to figure out what happened” tool. On the other hand, it makes the risk of leaving a laptop in the airport more risky than it already is. It increases the damage a successful malware deployment can do.

It also increases the amount of stuff that you potentially have to store for legal holds. Many industries have an obligation to hold on to certain things for extended periods of time. Recall is liable to include information that would fall under those regulatory holds so, congrats, now IT also has to implement an archival process for all of those across their entire fleet.

I guess that’s a long-winded way to say “depends on the company, but please for the love of all that is holy do not do personal stuff on your work’s laptop”.

I, personally, have no idea why anyone would trust Microsoft enough to keep this feature secure, and I’m pretty sure we’re going to see a stark uptick in attacks targeting this feature as it rolls out.

Likewise, it would not surprise me if at some point in the future an update to “send select metadata to Microsoft” pops up because the AI race is fueled by data and the allure of all that distributed data is strong.

I can see how this could be useful. I’ve frequently wanted to find something amorphous that I wasn’t able to readily find…. But the downsides far outweigh that benefit for me. I suggest you strongly weigh the risks vs. the benefits and don’t use this thing.

Apple Intelligence is also not great

Apple Intelligence in a rainbow color

Update June 17th, 2024 - Ed Zitron has a pretty great take on this where in he argues that the ChatGPT integration is a mere footnote and it doesn't seem like Apple thinks it's going to be useful.

Well here we are again. Yet another giant company has decided to put Generative AI front and center, embedding it inside the operating system. I'm starting to think this is the most cursed timeline possible (not really, but the simulation is getting real weird).

I'm going to get to that, but first I want to take a minute and talk about that WWDC keynote, because it had some other things that I liked and a bunch more "AI/Machine Learning" than they actually mentioned.

Look, right out of the gate, Apple announced a ton of things for each one of their operating systems. They did the thing that Apple always does: Refine some feature something else has had for years and announce it like it was both their idea and revolutionary. They also did something I thought was smart: When a feature is powered by machine learning, they didn't mention that and just talked about what the feature could do. Brilliant! There's a lot of relevant, applicable things that ML can do for you that doesn't sway into the generative category.

These very handy features include:

  • Categorizing your emails into tabs, a feature Gmail has had since 2013. This is almost certainly using some kind of classifier model to do the work. It'll also likely be a bit of a hot mess (but that's okay).

  • Showing you relevant bits of trivia about the movie or show you're watching, a feature Amazon has had since 2018ish (Prime Video X-Ray)?

  • Taking 2d Images and making them into "spatial photos" (giving them depth) using an ML algorithm (I dunno if there's anything comparable to this one)

  • Shake your head while wearing headphones to decline a call from your Gam Gam (using an ML algorithm to classify a nod vs a shake), a feature some other buds did in 2021.

  • Tapping your fingers to do actions on your watch, which again is using a classifier model to understand what gesture you just made.

Sidebar, did anyone else catch the dig at Google Chrome in the Safari announcements? They basically said "Safari is a browser where private mode is actually private" - which is an amazing throwback to this revelation.

By not actually saying anything about ML in those announcements, you instead focus on how things actually make your life easier rather than eating up the current hype.

Hands down, the star of the show is the Apple Calculator App for iPadOS. Why the hell the calculator has never been available on an iPad before has been the subject of many a blog post and article.

But the absolute coolest thing that was announced was the handwriting scratchpad. Simply write your equations out, and the app will do the math for you. It has variables! It can do algebra and trigonometry! It'll add up the items in a list you just wrote down.

You know how it's doing that? Optical Character Recognition (OCR) which is another ML Algorithm.

Relatedly, that's also how it's doing the "your handwriting is bad and we're gonna make it look less bad but still recognizably like your handwriting".

I was vibing with the keynote for the majority of the presentation. Useful features, minimal risk, finally a calculator. I was a bit disappointed that they DIDN'T FIX STAGE MANAGER, but otherwise it was a good start.

And then Craig walked out to tell us all about ..... Apple Intelligence, immediately triggering my gag reflex.

Apple Intelligence is the same thing as Recall

Permalink to “Apple Intelligence is the same thing as Recall”

...but from a company with a better track record on security.

First thing's first, I want to say that Apple did a good job on putting privacy and security first. It's clear that they want to position themselves as the "privacy alternative" to Microsoft, and they've done a reasonable job of that over the years. Not perfect, by any stretch, but reasonable. Keep that in mind as we have a little chat about this.

Apple Intelligence, according to Apple, is a bunch of models running primarily locally on your Apple device provided you've got a strong enough chip (more on that in a second). If the local models determine they can't do a task for you (unknown how it's making that decision, probably going to be on a per-feature basis), it'll farm that out to the cloud, but using "Private cloud compute". That's something Apple just cooked up and has a lot of info around how they plan to do it. They're even opening it up to 3rd party security researchers. Neat!

I'm just a security enthusiast, not a professional (though I do love a good HackFu), so I'm going to leave the "does this do what they say it does" to others. What I will mention is that Apple's reputation for security is miles better than Microsoft, so I suspect the general public is more inclined to believe them when they say something.

It's also going to do Recall-esque things across your entire device, because it has deep access to all the data on your device. It won't be taking screenshots every few seconds because it doesn't have to. It already has deep systems-level access to aggregate everything you're doing. They've been doing this for a while now, for example if you've seen someone send you a picture in iMessage you can also see that same picture showcased for you in Photos. They've been categorizing people, places, and things in your photos for a long while.

Now, because of the deep integration of all of those things, they can semantically search and gobble all of that up at once. Yay!

Look, Apple does have some goodwill left to give it the benefit of "we're not going to pillage this juicy training source and we pinkie promise it's more secure than Microsoft", but this is the same thing.

These local and cloud models will apparently do all sorts of great things:

  • Make Siri more useful! (maybe)

  • Summarize your emails!

  • Summarize and prioritize your boundless notifications!

  • Summarize your text messages!

  • Semantic search across your entire device (just like Recall!)

  • Semantic Search inside of videos! (Possibly getting the subject very wrong)

  • Write a bedtime story for your child!

  • Check your Grammar for you! (I'm not sure why you'd use an LLM for that as I've mentioned before)

    • RIP Grammarly?
  • Make an Emoji of your friend, without their consent, and send it to them!

  • Make a custom emoji of whatever you want!

  • ...Generate "art" locally, I guess?

Basically, exactly the same stuff that all the other LLM / Diffusion model companies want you to do smashed together with the idea behind Recall.

Let's take a moment and look at some of the examples they gave

Permalink to “Let's take a moment and look at some of the examples they gave”

This, for me is one of the key things that illustrates just how "we're grasping for ideas" the LLM/Diffusion crowd is.

Enhance your notes with hallucinatory diffusion

Permalink to “Enhance your notes with hallucinatory diffusion”

Notes about Indian Architecture with a nice sketch about said architecture

The one that really got me, and I thought was absolutely ridiculous was an example of a student learning about architecture. They'd drawn a fantastic sketch of a building that they were learning about. BUT WAIT! That image is just a sketch. What if we had a machine hallucinate a copy of it? The demo goes on to show how a generative model takes the sketch and makes it into a picture.

Who the hell wants to do that? First off, it was a great sketch, and didn't need to be "enhanced". Second, you're learning about something that presumably...exists. That there are real photos of. Why would you waste the electricity to make a brand new image, that may or may not look like what you wanted when you could...I don't know, find an actual photo of the thing. Wikipedia has one, even! I just... ugh.

It's one thing to want to generate an image of a fictional character that doesn't exist and a whole other thing to want to generate an image of things that absolutely do already exist.

I don't have a lot of energy to talk about this one, but if you thought memoji were weird, at least you had control over that. This time, because Apple Intelligence "knows your friends" you can now make a memoji for them whether they like it or not. Doing...who knows what? Can I type in whatever prompt I want?

I do not like this. Please do not send me any of these. I might block you.

Relatedly, you're going to be able to do contextual diffusion models right there to "really express what you're feeling". I honestly do not understand why anyone would want to do this. I don't get it. Please do not explain it to me. I don't want to know.

Anyhow, those are just two examples of "I don't think they've got a killer use case for this so we're going to guess at what people do with their phones". The real problem is the giant OpenAI shaped elephant in the room.

All the Privacy in the world stops when you send stuff to OpenAI

Permalink to “All the Privacy in the world stops when you send stuff to OpenAI”

The other "major announcement" that Apple made at WWDC was their partnership with OpenAI.

Somehow, which wasn't exactly clear during the keynote, if a local model determines that it can't do a thing for you - and the private cloud models can't either, it may prompt you to send your query to OpenAI. This might just be isolated to Siri, or it could be everywhere? I'm not sure.

It's not exactly a "groundbreaking" integration. It's basically the same thing as any other OpenAI API wrapper, but at the OS level (Yikes).

The major problem with this is you may trust Apple not to leak your data everywhere, but you sure as hell shouldn't extend the same trust to OpenAI. Apple assures us that OpenAI isn't retaining that data but OpenAI has been caught lying repeatedly. You cannot trust OpenAI to do what it says it's going to, and shipping an integration to OpenAI right in Apple operating systems is a major breach of trust.

I don't like this one bit.

I hope that there's a way to disable it.

Notably absent from Apple's announcements is "what happens if the thing hallucinates". All we got was a single line at the bottom of a screenshot saying to "Check important info for mistakes". Awesome. We're going to have LLMs summarize everything on your device, and we're not going to mention that it's liable to get those things wrong. Cool.

I expect, especially with locally running models, you're going to get more hallucinations, not less, but Apple seems to think the opposite. Guess we'll find out.

Well, here's a silver lining. I have an iPhone 14 which will not support Apple Intelligence (excellent). So I guess I'm just never going to buy a newer iPhone? My iPad and macbook do support these features so... I'm going to wait until I hear about how much of it I can opt out of before I update to the next version of the OS.

It's possible I'll be sitting on an old OS forever.

Or I'll go live in the woods I guess. I do not want any of this, and it's being foisted upon me because our tech overlords think it's what investors want. The enshittifcation will continue until morale improves.

Putting Bluefin on a Surface Go

Surface Go Showing the Bluefin Desktop

So, since Microsoft decided to shove AI into everything, destroying what little trust I had left in them in the process, I’ve been giving Linux another go! Now, disclaimers first, I’m still primarily in the Apple ecosystem. All my Cthonic Studios work is done on Macs (and this is partly for software reasons, which I’ll get into in a future post), and I have some Windows Gaming Handhelds (also more on that at the end) but I’m slowly converting some of my tech over.

Today I want to talk a bit about putting a Linux distribution on a Surface Go 1, which is something I wish I’d had the ability to read more about when starting this project, but before we get there I need to do a bit of preamble.

Back to Linux, a surprisingly pleasant journey

Permalink to “Back to Linux, a surprisingly pleasant journey”

A couple of weeks ago, I put Bazzite on my Framework 13”, and it’s been absolutely stellar. I’ve had zero major issues (and only a small number of minor issues) with that. Games still work! Doing development work still works great thanks to magical containerization! The wifi drivers just work! all in all it just works. It was a wonderful experience.

And all of this is thanks, under the hood, to Fedora.

Bazzite and the other distributions is based on Fedora Atomic Desktops, which are really neat. They’re basically the latest iteration of immutable desktops, which use various techniques to ensure that stuff in user-space remains isolated from the bits that make the computer go “brrrrr” which makes it easier to rollback or make changes to the core OS more safely.

One of the major barriers for me using Linux on the desktop for the longest time has been the propensity of updates to completely bork the system, making it unusable and taking upwards of hours to recover. Fedora Atomic Desktops promises to remove this problem by keeping updates contained and allowing you to boot back into previous versions of the OS after updates. You can even rebase your own version against another version to get yourself back to a known-good state in minutes and pin yourself there until the problem is corrected.

This is amazing. Universal Blue (the people behind Bazzite) also have a handy support document which covers how that functionally works. I’ve done a few updates on the Framework and so far nothing has exploded, so I haven’t had to resort to any of this yet, but just knowing it’s there and how easy it is to roll back is a comfort.

So, I’ve had this Surface Go 1 since it was first released and it’s gone through a series of being wiped about 4 times. Most recently I put Windows 10 back on it and it was…well… painful. The whole thing is super slow, and while the AI nonsense hasn’t made it everywhere, that damned Copilot button did get popped in. I don’t like this.

There’s no reason to get a Surface Go 1 in 2024. I’ve had one sitting around that I’m saving from eWaste and it seemed like a fun project.

So, it was time to try something different. I was so impressed with Bazzite, I wanted to see what it’d take to get Linux running on the Surface.

Here’s where the research started to get a bit concerning.

Running Linux on a Surface device requires a specialized Kernel

Permalink to “Running Linux on a Surface device requires a specialized Kernel”

Right, so the very first thing you’ll run into when wanting to run Linux on a Surface is that because the hardware is specialized you’ll need a specialized kernel. Fortunately, the folks over at linux-surface have done all of that workand even have a well documented way to get it working. That said, if you start reading the documentation, there are several steps you have to take post-installation to get the kernel installed.

The easiest way, of course, is through the package manager but for some surface devices you’ll need an ethernet port because the Wifi chip isn’t going to work out of the box.

But, I’m lazy. I don’t want to do that. Universal Blue to the rescue.

Universal Blue Images have linux-surface kernels built in

Permalink to “Universal Blue Images have linux-surface kernels built in”

That’s right kids, Bazzite and its peer systems (Bluefin, Aurora, uCore) have an image you can download which includes the linux-surface kernel in it.

Since I figured the Surface Go wasn’t going to be giving me a lot of gaming mileage in 2024, I went with Bluefin with the developer tools not preinstalled.

Sidebar, near as I can tell Bluefin and Aurora are the same except for Aurora uses KDE and Bluefin uses Gnome as the desktop manager.

This turned out to work great. I had zero problems wiping the entire drive, installing Bluefin, and then running it, though I did run into a couple of issues.

Surface Go showing a text editor side-by-side

This was an extremely straightforward process, so I’m just going to use a numbered list:

⚠️ Now, you might notice that below I don’t re-enable Secure Boot. This is because I didn’t update the BIOS before starting this process and older versions of the UEFI menu do not allow you to boot 3rd party systems with secure boot. I’ve not tried fixing this yet via linux, but I will update this post if I get around to it. So if you’re coming behind me, update your BIOS to the latest before doing this using Microsoft’s update tools.

  1. Download the surface image from the Bluefin site. Follow the prompts and they’ll give you the right options.
  2. Flash the image to a USB stick using the Fedora Media Writer (or using Balena Etcher)
  3. Reboot the Surface, and hold Volume+ while pressing the power button. This will boot you into the UEFI menu.
  4. Disable Secure Boot, and reorder the boot order to have external media be up top (getting the older Surface to show me the boot menu was difficult).
  5. Plug in your new USB drive and reboot (I used a USB-C stick, and did the whole install on battery).
  6. Boot to the linpus-lite drive and follow the prompts to install Bluefin.
  7. It will prompt you on the next reboot to Enroll the MOK for Bluefin to allow you to reenable secure boot.
  8. Profit!

After doing all of that I had a fully-functional Surface Go running Linux! Here are the things that work which surprised me immediately:

  1. The keyboard cover! (though it does take several seconds to recognize it on a fresh boot which is weird).
  2. The Pen is recognized as an input device! Pressure Sensitivity actually works in Krita!
  3. USB plug-and-play
  4. Trackpad gestures!
  5. Screen Rotation!
  6. DisplayLink Drivers for my external dock!

Surface go Docked to an external 2k monitor

All in all, everything again just works ™️ which was certainly not my experience for Linux on the desktop the last time I was daily driving it.

This thing isn't going to win any speed contests any time soon. Earlier, I was watching a youtube video while trying to multitask on it to see what it'd do, and while it behaved admirably given its aging hardware, it still stuttered when displaying the second display at 2k, playing video, opening discord, and trying to open LibreOffice.

Once everything was running, it wasn't a bad experience (certainly better than windows), but it was still kinda sluggish.

The other thing that I had thought / hoped we'd solved is the blurry bits on a high DPI display with Electron / Chromium apps. The tl;dr here is you have to enable the Ozone settings and set it to use ozone as the renderer in each of these apps flags: --enable-features=UseOzonePlatform --ozone-platform=wayland which will fix the display to work properly.

Update: August 2nd, 2024: Turns out this is a Gnome fractional scaling problem. I've installed Aurora (which is the same thing but with KDE) and the problem goes away.

One final bit: you need to remember to suspend it manually if you don't want the battery to drain while you're not using it. I suspended it and left it overnight (for around 10 hours before I checked) and it'd drained about 10% in suspend, which was pretty good all things considered.

But yeah, otherwise everything worked pretty well!

So, I also have a Aya Neo Air which I’m barely using and guess what “mostly works” on it? ChimeraOS. So, I’m going to tinker around with getting that installed and working on it in the nearish future and see how it goes. I’m not 100% sure I want to do this since I’ll lose some of the Aya Neo niceness (for example, the TDP switcher is broken in Chimera so you have to use a different app).

I can’t do this on the Ally as no distribution will ever properly work with that proprietary connector for the XG Mobile so I’m going to have to keep playing cat-and-mouse with Windows on that one, but…yeah.

Anyhow! Hopefully I’ll get to report back with good news next time.

Taking /e/OS for a Test Drive

Pixel 6 Running /e/OS on my desk

Right, so here we are again, just another step in my journey to find a way to opt out of all the “shove generative AI into everything” trend in my personal life. In my last post on the Apple Intelligence announcement, I ended on a note of “well, I guess I’ll just not upgrade”. That’s a viable strategy for a while (which I intend to make use of for as long as I can), but in the long term it might not be viable.

Look, I’m not totally opposed to running an LLM or something on device, but I want it to be on my terms, not baked into every interaction with the operating system. Apple seems best poised to allow that kind of behavior (which is funny given their track record on customizations), while Google and Microsoft are hell bent on shoving this everywhere in search of the next big thing.

Thus, I spent a week running an experiment running one of the many projects based on the Android Open Source Project (AOSP), the open source part of the Android project sans the Google Extras. I figure the LLM stuff won’t make it into the open source side of stuff maybe ever since Google sees it as a moneymaker (or do they?), so it seems like a reasonably safe bet for a while.

In order to simulate what my life would be like in a world where I try to get away from the big players while still holding a smart phone (a dumbphone or Lightphone are always options), I set some rules for myself.

  1. No signing into Google services on the phone. They shall not touch the datas.

  2. No signing into Apple services either. It’s not as big a problem, but I do use Apple Music all the time.

  3. Stay as close to the OS’s preferences as possible.

That’s pretty much it, but I did a bit of a bonus goal while I was doing this. I paired a CMF Watch Pro and used the CMF Buds Pro as my daily drivers for the same time period. I wanted to go with a bit of a “budget premium” experience to challenge myself on the “do I really need the best stuff”

As an aside, my primary watch is an Apple Watch 6 and I see no reason to upgrade.

The gear list for this experiment is:

  • Pixel 6 Pro (Used, from Backmarket)

  • CMF Watch Pro (Amazon)

  • CMF Buds Pro (Amazon)

There are a lot of different flavors of AOSP projects out there for various needs. Lineage OS is probably the most common one out there, and Graphene OS is the extreme privacy / security focused version, but I went with /e/OS because I was already familiar with it and it offers a very interesting sync capability (which I’ll get to at the end of this post).

/e/OS is itself a Lineage fork, by the by

Now, I could have gone off the board to a Linux based phone distro (and I may give that a go in the future) but that’s a big jump I didn’t want to take - and I wanted to try to preserve the “app ecosystem” that I’m used to. Which is a big advantage of /e/OS: it ships with MicroG (which they maintain) and an app store that can pull APKs from the Google Play store anonymously (thus fulfilling rule #1 but still taking advantage of the Play store ecosystem). Jumping to only using progressive web apps (PWAs) and open source apps was a bit too far of a stretch for my daily driving.

Now that I’d picked an OS, I needed to pick a phone that would both be compatible and ideally wouldn’t create more eWaste in the process. My original plan had been to use an old Galaxy Note 10 that we had lying around, but unfortunately Samsung US phones have locked bootloaders that are a pain to unlock. So, I went on Backmarket and found a Pixel 6 Pro Unlocked, and in good condition. I did this for both price and compatibility reasons, before you set out on this endeavor yourself, you’ll want to read the supported device list for your given OS carefully.

Unlocked from Google is an important distinction here. The ones the Carriers sell may not be able to unlock the bootloader as well, leaving you stuck with stock android.

The phone arrived safely and was in pretty good condition. There were a few scuffs, but nothing a thin screen protector couldn’t hide (sparing me the madness of noticing it all the time).

Alright! Let’s go!

/e/OS also has an “Easy installer” that supports some newer devices, which is very fancy and easy. Unfortunately, a Pixel 6 Pro is not one of the supported devices.

Okay, so this part is very straightforward, but requires a few steps and you need a separate computer to do it. That guide is here. At a high level, the process is:

  1. Download 2 files from the /e/OS site

  2. Ensure Android SDK tools are installed on your computer.

  3. Use adb / fastboot to reboot into recovery mode.

  4. Use fastboot to unlock the bootloader.

  5. Flash the files from step 1 onto the phone.

  6. Reboot and enjoy your new OS.

Okay, we need to take a quick second and talk about unlocking the bootloader for those who are unfamiliar. The bootloader tells your phone how to load Android, and it is locked to an official version of Android so that you can be sure that the OS running on your phone is what you and the manufacturer expect it to be. It’s one of the things that prevent hackers with physical access to the device from doing nefarious things.

Now, unlocking the bootloader is both what allows you to flash a different OS onto the device, but it also breaks that integrity guarantee, meaning if a threat actor grabs your phone and has a few minutes, they can do some shenanigans like pull all the data off of the phone.

Some devices support relocking the bootloader. The Pixel 6 Pro is not one of those devices. This can still be fiddly even if it is supported.

/e/OS takes a defensive step against this by encrypting the data at rest, encrypted with your password (or pin) so that if that were to happen, at least it’s behind an encryption wall. Now, they could also potentially load malware onto it, which then also hijacks your password and conceivably sends it somewhere, but those are the risks we take.

At its core, /e/OS is just the stock Android 13 experience without the google services built in. Instead, it ships with MicroG, which is an open source implementation of most of the Google APIs.

What this means is that most Android apps, even the ones that rely on Google Play services will still work. (SafetyNet also works in some contexts, though I don’t have any apps that are using it). You can control / enable / disable MicroG at your leisure, but it is necessary for push notifications to work. In particular, Push notifications must still register with Google services because that’s just how push notifications have to work - there are no alternative push servers. There are some apps that allow polling for notifications, but if you want them in realtime, this is your only option (that I am aware of). These are still anonymous as documented here.

The default Launcher (Bliss) is a lot like older iOS before the app drawer existed - namely, there is no app drawer. All installed apps just appear on the home screen. Also, you can’t place widgets anywhere except the leftmost screen. I’m not the biggest fan of this experience, but I stuck with it for this experiment. If I were to do this long term, I’m liable to install an alternative launcher (I like Nova).

A stock Android phone without installing apps will not be very useful unless you’re going for a minimalist kind of thing, but that’s not me, I want apps. You do have the option of only installing open source apps, but that’s also not going to get me to a fully functional app (Look, I use Discord a lot).

App Lounge has a very handy feature so that I would not break my first rule: You can use an anonymous Google account to download apps from the Play Store. This requires you to additionally trust the /e/foundation to not tamper with the apps as they’re being installed, but the app is open sourced so it can be audited and it wouldn’t behoove the e foundation to slip in some back doors. Anyway, using one of any number of anonymous accounts they maintain, you can essentially side load apps from the official play store, which means I can install Discord without too much trouble.

The other neat feature they’ve baked in is the “Advanced Privacy” which essentially blocks ad trackers at the OS level and gives you a handy report on every tracker it’s blocked and which app the request originated from. This includes stuff like the usual Firebase analytics, all the way down to some less common ones like Qualtrics.

I’m pretty sure they’re doing this purely on the domain of the request and using a block list to make the determination, which means some stuff is probably slipping by, but it’s pretty fun to see that DoorDash is extremely leaky.

Likewise, the Aurora store gives most apps a privacy score out of 10. You’ll see a lot of them with a 0/10 on the privacy list, because of all the tracking.

Let’s talk about Murena Cloud (formerly /e/Cloud)

Permalink to “Let’s talk about Murena Cloud (formerly /e/Cloud)”

So one of the other really neat things about /e/OS that I wanted to mention is their cloud service, which they’ve got an integration directly in the OS which allows you to sync files, notes, contacts, emails and the like as you would if you were heavily in this ecosystem. Murena.io hosts this cloud service, which gives you 1 GB of storage for free and charges a reasonable fee for upgrades.

But we’re not in this to simply trade one cloud provider for another, oh no. The coolest part about the cloud software is that it’s essentially a specialized fork of Nextcloud, and it’s open sourced. That means, theoretically, you can run the whole set of cloud services on a server you control, deeply integrated with your mobile operating system. That’s pretty neat.

Now, you can already do this with vanilla Nextcloud and the various Nextcloud apps, but you’ll lack the built in integration. It’s not a big hassle to just use Nextcloud, but it is really appealing that built-in integration supports private servers.

The other thing that is cool is that murena.io also just works as a Nextcloud server, so you can do things like use the Nextcloud Feeds app to sync up with RSS feeds you put up in murena.io. Or use the Nextcloud files app to sync your files if you really want to. Lots of flexibility there to host or not host what you want.

Now that the experiment is over and I’m back in the corporate embrace of Apple, what did I like about the experience?

Overall, I was really impressed at how smoothly everything went… for the most part. I could install all the apps I needed, and the PWA support for other major apps (like Starbucks) meant I didn’t need to install as many apps as I have before on other Android phones.

I was also blissfully free of the “TRY GEMINI!” pop-ups that have started appearing on other Android phones (like my Motorola G5 Power) or whatever Samsung is calling theirs (which also just appeared on my wife’s Galaxy S22).

I was concerned that I wouldn’t be able to install an eSIM and get it working, but I was also pleasantly surprised that worked without any issues as well.

The only major issue I ran into whilst doing this experiment is that I make heavy use of Apple pay on the Apple watch, specifically to store my transit card. Tapping and getting on transit is an essential part of my day, which was sorely lacking on the /e/OS build. Payments are probably never going to work on the phone directly (because of how payment industry infrastructure works and requires direct partnerships).

There are some workarounds. I decided to see if I could get my old Galaxy Watch Active 2 to pair and verify a card via Samsung Pay on the watch, and I’m happy to report that after some wrangling, I got it working. The only major hurdle is that I had to install the Edge browser and make it the default so that Samsung would let me log into the Galaxy Wear app.

Now, this is a Tizen based watch, not a WearOS watch, so I have no idea how a WearOS watch is going to function, but reports on the forums don’t look very promising. If I wind up trying it out, I’ll write a follow up post.

The forums have had some success with Garmin Watch / Garmin Pay, but the banks that Garmin Pay supports are pretty limited in the US.

Either way, I got the Watch Active 2 to verify a card, and I was able to tap and pay at a Petco shortly after verifying it. So, uh, success!

Overall, I think this was a pretty successful experiment, and should I be unhappy with my ability to disable OpenAI integration in iOS 18 I can at least switch my phone if I need to - but given that I can just stay on 17 for an indefinite period of time I’m unlikely to make the jump just yet.

I'll update this post when I've got the review of the CMF watch up with a link to it!

What the heck is going on over at Proton?

tl;dr I’m no longer recommending Proton services to anyone and have moved off to other services.

Right, so up until recently I was a happy Proton customer at the annual unlimited tier. They were missing some convenience features and the integration of a document editor into the Drive product was pretty great.

That is, until they decided to shove an LLM into their core product and upend their security model and break some trust by the way they rolled it out. Let’s take a quick look at their timeline / speedrun of adding LLM features (and Crypto wallet, but we’ll get there in a sec).

  1. June 5th, 2024 - Proton releases the results of their 2024 community survey
  2. June 17th, 2024 - Proton announces transitioning to a non-profit structure
  3. July 8th, 2024 - Eamonn Maguire posts about building "privacy protecting AI" - Signaling that they were "thinking about" the problem (but clearly had already built the thing).
  4. July 18th, 2024 - Proton releases Proton Scribe - this is the LLM product.
  5. July 24th, 2024 - Proton releases a Bitcoin Wallet

Now that we’ve established the timeline, let’s break down the path from Survey to “releasing products that no one asked for”.

So I took this survey, and I’m now kicking myself for not screenshotting the questions, because I have a lot to say about shitty survey design.

I’m going to do some conjecture and speculation here, but I think that the survey was designed when they already had the Scribe product already well into development and that the questions were designed to make it seem like the product they were already developing was by popular demand. Like, you don’t develop a “privacy preserving” (it’s not) LLM in 1.5 months. They were already building this thing.

Let’s talk about those results. They provide this graph in their survey results post:

Chart showing that 29% of users want a "writing assistant"

My recollection of this question is that it was a multiple-choice, but not stack-ranked question. I’m not sure if that’s how it actually was, just how I remember it.

Regardless, I want to point out 2 things:

  1. The LLM answer in there doesn’t actually mention an LLM. It mentions a “writing assistant”. There are tons of things that do not use Large Language Models to do writing assistants, like checking grammar and spelling. The way they worded that answer was extremely misleadging.
  2. Only 29% of respondents said they wanted it.

So let’s move on to the second point they try to make in this survey that absolutely does not say what they say it does.

Chart showing percentage of users who "have used" AI

And here’s their analysis of that data:

Generative AI is one of the most significant developments in recent history, and it is supposed to lead to incredible gains in productivity. As more and more AI assistants come online, we asked the Proton community what they thought of these tools. Around 42% of respondents use an AI service regularly (at least once a month), and another 18% have never tried AI but are interested in it.

I just. sigh.

  1. The survey asked do you use Generative AI. It doesn’t ask why. It doesn’t ask if you find it useful. It doesn’t ask if you WANT IT IN YOUR FUCKING EMAIL CLIENT. This tells you nothing!
  2. “At least once a month” is not very much Generative AI usage.
  3. That’s still under half of your user base!

This question is poorly designed. I can’t tell if it’s poorly designed as an excuse to interpret the results the way Proton apparently wanted to, or if it was a genuine “surveys are hard to design” take, but asking “do you use Generative AI” without also asking “Do you use Generative AI for spicy roleplay” and “do you want us to put an LLM into your Email client” is ridiculous.

So, I don’t think this survey is a killer “Our users want this” result.

Testing the Waters - "What if we made a privacy focused LLM?"

Permalink to “Testing the Waters - "What if we made a privacy focused LLM?"”

On July 8th, about a month after producing the survey results, Eamonn Maguire posts about building "privacy protecting AI". This post reads to me like a love letter to Generative AI and was super suspicious to me. I posted on Mastodon in response to this post with the assertion that I'd immediately move off of Proton if they went forward with this.

Introducing an LLM, for our “business users”

Permalink to “Introducing an LLM, for our “business users””

10 days later, Proton releases Proton Scribe. Since you can't implement something complicated like this in 10 days... they already had this ready to go.

In their announcement post they make the following assertion:

In our 2024 community survey, more than 75% of Proton’s business users said they are interested in generative AI tools, but most were also concerned about a lack of data protections. Scribe was designed to be a secure alternative.

Right. Okay, I don’t think that’s what those survey results actually say (since you didn’t ask “do you want a secure alternative if we build one”).

When the backlash to this feature started over on Mastodon, they had the following response:

Screenshot of a Mastodon response

This excuse is flatly ridiculous. “We thought ‘hey they’re gonna do it anyway, let’s do it so they can be safe’”, is what that amounts to. That’s not a smart way to roll out features - and I don’t actually buy it.

Meanwhile, over on a Pivot to AI blog post on this subject, Amy and David point out the following:

Proton’s descriptions of Scribe are vague and waffly about their threat model. Your prompt — that is, the email you’re writing — is kept in plain text on their server, unlike emails you’ve sent or received, which are secure at rest. Proton promises they don’t log the prompts — but services like Apple, which many Proton users were trying to get away from, make only the same level of promise.

Proton then goes on to point out that the Model can run locally…. but only on Chrome, and only on systems with high enough system specifications. They went on to respond to this post with the following rebuttal on Mastodon:

Screenshot in response to a Mastodon post

I want to break down some of these things.

The feature is not “opt-in” if you’re on affected plans. It shows you this dialog:

Proton Scribe showing the opt-out dialog

Notice how that dialog doesn’t have a “no I don’t want this” option. To turn it off, you have to go into the settings. That’s an opt-out feature. Not opt-in.

The second thing is that there’s no such thing as an “Open source” model. The weights may be open, but Mistral has been super cagey about where they got the training data (also known as, they got it the same place all the other companies did - the internet, without permission).

Finally, if you or someone you’re chatting with decides to use the “zero logs server” the text is sent to them in the clear…. which I think very much does break their zero-knowledge model.

You’ve totally broken your privacy model for your users, willfully. For an LLM.

Ugh.

But that’s not the most ridiculous thing about this whole saga.

Literally six days later, Proton announces they’re launching a Bitcoin wallet of all things!

Pivot to AI covers this too and points this out:

If Proton was taking privacy seriously they’d have used Monero or Zcash — two cryptos that use zero-knowledge proofs to make transactions untraceable through the blockchain data trail. At least these have a use case, even if it’s buying Russian research chemicals off the darknets.

Yeah, exactly that. Bitcoin isn’t private, and no amount of you providing a non-custodial wallet is going to change that.

Here’s the security model. This is a reasonable attempt at the functionality, but it doesn’t make the basic idea any less dumb. Also, they recommend you use a bitcoin mixer so the Feds can bust you for money laundering too when they match you to your bitcoin address.

… That’s bad.

I really wish I knew. The timeline makes it clear to me that they had several potentially very unpopular features already in development and then timed the release with their developer survey. I suspect they were hoping to spin the results as a “you asked for it and guess what we delivered in record time!”.

The spaces I hang around in are very skeptical of both Generative AI and Crypto, and I think that Proton’s core individual customer is too. It feels like they’re speedrunning a “alienate our core customer base” but I honestly don’t know what they’re doing here.

It’s possible that their business users really do want this thing and have been asking for it for a long time. It’s also possible that the people running the show are adding features for a future investment or a sale. That doesn’t really jive with moving to a non-profit.

It’s also possible that they’ve always loved LLMs and Crypto and just happy to shove it into everything.

But, it doesn’t functionally matter for me. They’ve broken my trust, and that was the primary currency and reason I’d be willing to shovel money towards them. That takes a while to build. It takes an instant to destroy.

I think about that a lot. Oh well.

ChimeraOS on an Aya Neo Air

ChimeraOS Logo

This is a follow-up post to  Putting Bluefin on a Surface Go 

In my quest to put Linux on absolutely everything that used to run Windows that I can, I’ve already moved my Framework 13 over to Bazzite, my Surface Go to Aurora, and now the Aya Neo Air over to ChimeraOS. Things are going well-ish.

So, for the Air, there are a couple of options - I could have tried to put Bazzite on it but so far no one on the forums has tried it on the original Air model so I was a bit wary of that whole situation. On the other hand putting ChimeraOS on the original Air is a well-trod path and almost everything works out of the box according to the hardware compatibility page.

Interestingly that page marks “eGPU” support as “?” - it shouldn’t even be listed as the Air doesn’t have a USB-4 port or oculink or anything so outside of a case hack you’re not running an eGPU on it.

Likewise, there are several Youtube videos out there documenting the process and some of the pitfalls. I followed this one:

Setting the Air up to install ChimeraOS was actually pretty easy, once I got past actually getting it to boot from a USB stick. There are a couple of guides that state you can get into the boot menu via various tricks with the power + volume rocker, or via the normal keyboard shortcuts. The only thing that actually worked for me was hitting the delete key to get into the bios and manually changing the boot order.

As a result, I had to reboot like 6 times. The vibration motors buzzed the entire time. It was… not quite maddening, but it was starting to get on my nerves throughout the process. Buzz. Buzz. Buzzzzzzzz…..

But, past that, Chimera’s installer is just a couple of options and it will process the whole thing. At the time of this writing, it is not possible to dual boot ChimeraOS and Windows (but this is possible on Bazzite), so it’ll completely overwrite your Windows partition. I did make a backup of my product key before doing the wipe but it’s unlikely I’ll return to Windows on it.

Once the installation was complete, it booted straight into the SteamOS interface without any issue. Login was seamless and within a couple of minutes I was installing games and ready to go.

One issue that you’ll want to be aware of is that TDP control (at the moment) doesn’t work via the standard Steam controls, but it does work fine from a Decky plugin and a few other methods (I used SimpleDeckyTDP and it works great).

Overall, performance is basically identical to what I was getting on Windows without the hassle of Windows causing problems every time I boot. The whole system caps out at a 15W TDP, so you’re not going to get great performance on AAA Games - but here’s a sample of games I’ve tried on it at the full 15W:

  • Dead by Daylight - 720p - FSR on, 40+ FPS stable
  • Skyrim - 1080p - Medium Graphics preset, 60 FPS stable in Dungeons, 45 in the overworld
  • Neon White - 1080p - High Graphics, 60 FPS stable
  • FFXIV - 720p - Laptop (Standard), 40+ FPS, mostly stable unless you’re in Limsa or Tuliyollal.
  • Shin Megami Tensei III HD Remaster - 1080p - Standard Graphics, 30 FPS stable (SMT3 is locked to 30 FPS on this port)
  • Monster Hunter Rise - 1080p - Low Graphics, 60 FPS Stable
  • Resident Evil 2 (2019) - 720p - FSR 1.0 Quality

Which is to say, I don’t notice any difference between Chimera and Windows.

Here’s a visual representation of some of those:

I’ve only had a couple of problems with ChimeraOS since I installed it, the biggest of which is that sometimes it refuses to awake from sleep and I have to hold the power button down for an egregiously long time to get it to boot.

It’s happened twice where I was concerned that the console had given up the ghost. Fortunately, it hadn’t, but it’s been enough to trick me both times into a mini-heart-attack.

The other issue is you can’t really control the LEDs under the sticks (I think there is a way to do it while the console is on, but I’ve not futzed with it too much yet).

It’s been a fine experience, and I’m enjoying it. Way more than I was enjoying Windows on it. Would I do it again? Yes. Would I buy the Air in 2024 just to put ChimeraOS on it? No. Absolutely not. The battery life is terrible and there are better options.

Self Driving Car Update

About 10 years ago, I got up on a stage to give a lightning talk about self-driving cars. If you want to relive that person I was, it’s still up on Youtube (and my talks page, so I can’t be accused of hiding stuff). Here’s an embed:

Now, the idealistic person I was in that talk made a fair number of points that I still generally believe:

  • Dedicating a ton of space to cars that are not in active use rather than human-centric infrastructure is bad.
  • We should, as a society, work to reduce traffic deaths - especially those involving a car and a not-car (pedestrian, bicyclist, and the like).
  • Driving is a chore, and at the time I sorta resented living in a place where it was an abject requirement, both for me and folks who really would have been better served by on-demand transit.

At the time, I was pretty bought into the hype that we were making great strides in self-driving technology. I even at some point after that talk repeated the talking point of “we’re just five years away” (because we’re always just five years away). A lot of that idealism was born of a few different things: Tech was still a generally fascinating and exciting area of endeavor for me and I was still very optimistic that the industry generally wanted to make people’s lives better. Likewise, I lived in a place with horrendous public transit. Very few bus lines, 1 hour wait times, just generally unusable. Cars are a necessity there - and cars that can park themselves or be summoned on demand seemed like the best option.

I do want to give Oklahoma City a bit of credit here, their transit system has gotten a lot better in those 10 years - still not where I’d expect a major city to be, but it has markedly improved.

Likewise, there were a lot of demos being shown at the time faking cars abilities to drive themselves (Tesla, for example, in 2016 a couple of years after my talk - at the time it was the Google Incubator precursor to Waymo and Cruise doing demos).

I don’t really think these things anymore - and part of that is a change of scenery. I now live in a place and have visited places that have great public transit regularly. While trips can still take a lot longer than a car would, it’s not nearly as big of a burden as it used to be.

Sidebar: We do still have plenty of work to do, as always.

One of the things that I dismissed as impossible (because of my environment) was that we could build effective transit across the country and in places where perhaps the political will wasn’t very present has been replaced with a sense of displacement. The displacement is the persistent need from these automation companies to convince cities not to invest in other modes of transportation because why would you if the automation revolution is right around the corner. All that investment would be "wasted" (it wouldn't), so it behooves the automation companies to continue to say things are "just a few years away" for at least a decade, on repeat.

And like, progress is being made. Waymo is running a relatively-sanctioned experiment in a couple of amenable cities. Tesla is running an unsanctioned experiment with every driver who pays them more to do so (which has resulted in plenty of fatal crashes). But hey, these cars are running all on their own... or are they just offshoring the driving to remote workers? Hard to say for sure when the AI companies are constantly faking their metrics and demos. Long story short, they have lost the benefit of the doubt. They perhaps should never have had it in the first place, but it was a simpler time. Or at least it was (to me) without another decade of being in the tech industry.

Instead, I have personally doubled down on building out robust regional transit and investing in high speed cross-country and regional rail to cover our moving-people-from-one-place-to-another needs. For freight, I think we should focus on making working conditions for truckers the best they can be and maybe accept slightly longer supply chains (I think that we're going to have to deal with that regardless, we've already experienced several "historic" disruptions all in a row and climate change is only going to make it worse).

We need more bike friendly cities and more well connected suburbs. We should invest in transit at scale, it's a proven and time-tested thing.

Could we do that and create autonomous vehicles at the same time? Sure, yeah. However it's going to be very difficult to do in a world where politically we prioritize one thing over the other and defer to the private sector on too many things. If we were to do that, I'd also want something akin to a smart grid where the cars could know where each other are at all times, because otherwise you get scenes like this:

Endless cars, endlessly unaware that the other cars are also on the same system.

But such a system would also take us even further down to a dystopian surveillance state where your non-smart car is also tracked or you're encouraged to install a system to cover the tracking. Yeahhhh, maybe not great.

Anyway, in a world where we can't have both, I'd much rather have the good transit than the "maybe it works maybe it doesn't" promise of "just give us another 5 years and we totally go it, trust us".

So here's another thing that's a little bonkers about self-driving cars, invariably someone will say "Try it! It works great." And yeah, it just might! For you. For that one drive. What about when it holds up traffic? What if it breaks a traffic law and there's no one to ticket?

What matters is "does it work at scale" and "did we have to kill anyone in the process"? Waymo is probably being the most socially responsible one here, it's getting permission before deploying its fleet in pilot cities. Meanwhile Tesla's out there running an experiment on every road everywhere but prioritizing its CEO's drive.

Now, I imagine you saying "it just needs to be better than humans"! Sure! But by what metric? Humans know how to navigate construction without seeing that in their training data a whole bunch of times. And by who's account? If the data's all being held by the company and not being audited independently... who can be sure it's actually safer?

All of these metrics are currently vibes based. "It will get better! It just needs more data" is a vibe. "My Waymo ride was awesome" is a vibe. We're rolling out a consequential technology based on vibes and just letting huge US companies cop out with a "trust me bro".

But, the NHTSA is investigating currently as I write this, so we'll see what they find.

No. I hate being lied to about progress. I hate technology that isn't centered in our needs. I hate tech that's just being used to enrich a few giant companies.

That's why I don't think we're going to see large-scale self-driving any time soon. And that's okay.

Putting Manjaro on a 2015 Macbook Air

Manjaro on a MBA 11

As the world hurtles more towards shoving Generative AI into absolutely everything to make investors happy, I've been looking more into hardware sustainability and the viability of running things for longer (especially if they don't try to bake generative AI right into the operating system).

Something that's been chewing at the back of my brain for a while is the 11" Macbook Air. The form factor on that tiny laptop is awesome for how portable it is, but it was underpowered at the time for MacOS so I ditched mine.

Nearly a decade later, I keep thinking about it, it really was one of the best little not-a-netbook spiritual successors to those netbooks of yore. I kept thinking about how it might be nice to have a pretty portable writing machine for someone and I was curious how well it would work with that. So I've been keeping an eye on my local FreeGeek eBay store until I saw one in good condition, and snapped it up (for just over $100).

I also have an iPad which I use for this writing use case and I'll likely continue to do so, this is partially a fun thought exercise.

SpecValue
ModelMacbook Air 11" (7,1)
RAM4GB
CPU1.6GHz dual-core Intel Core i5
Battery Capacity38 Wh (31 Wh max on mine, 80% original capacity)
Storage128 GB

The 2015 MBA can only be upgraded to Monterey at the latest, and that's what came on the one I bought. I did boot into that and play around for a while and the experience was "okay". Unfortunately, running Monteray on something with only 4 GB of RAM is a bit of a rough go. It managed to browse the web okay but it was already a little sluggish just poking around.

Enter Manjaro!

To the former question, that's mostly because I can reliably run linux on a potato, and I was sure any number of distros would run on a nearly decade-old machine with strong memory constraints well.

The main issue was going to be drivers, older Apple hardware is usually fairly well supported but there are a bunch of little quirks that you have to deal with. So I set out to find a distro that was going to "mostly" work out of the box.

That's why I picked Manjaro with Xfce. The Xfce window manager is extremely lightweight (while still looking pretty great) and Manjaro is an Arch distro with pretty solid support for this particular model of Macbook Air. Installing it was a breeze, simply creating a bootable USB drive and we were on our way. Now, I expected to have to deal with Wifi driver issues after I installed Manjaro, because during the install wifi was not working. However, once I'd booted into it for the first time the wifi drivers just worked. That was a pleasant surprise. Wifi drivers have been my recurring Linux nightmare since the early 2000s.

After that, it all "just worked" with some minor exceptions, which were:

This is perhaps a huge deal for someone who wants to use this thing for video conferences, but it's not a big problem for me.

Luckily the Arch Wiki has a whole page of troubleshooting steps, webcam one of them. I'll eventually go through the steps and report back, but others have gotten it working with this particular model so I'm not worried about that.

By default, the screen's color profile is "muted" and off. Fortunately, Arch Wiki to the rescue again. The installation page has a reference to which color profiles to use.

This one shocked me the most, I use Emojis a lot on my self-hosted Outline instance. It was very apparent when I looked at it that they were missing.

Happily, this is also readily fixable via a single command: pamac install noto-fonts-emoji

Yeah, like everything else there's a reference on the Arch wiki. I've not tried this yet, but I'm also pretty confident it'll work. It doesn't take very long to boot the system cold, so it's not really been a problem.

The keyboard is still really solid and pleasant to type on. It predates all of the butterfly keyboard problems and I'm honestly a little surprised that there are no issues on this used model that I purchased.

The form factor is hard to beat. The extra length means it's super comfortable to work on, with lots of space to place your palms while typing. It's also got a full function row which is always appreciated (I'm looking at you, touchbar mac). Unlike netbooks, it's actually really easy to clack away on without feeling too cramped.

Surface Go 1

iPad Pro M1

Macbook Air M2

iPad Pro M4

I've gotta say I'm a huge USB-C fan, it makes life a lot easier if you only need a single cable for everything from external displays, charging, data transfer and the like. I'd forgotten how annoying it is to have to carry an entire magsafe charger and regular USB-A cables and maybe a thunderbolt 2 cable (not that this thing's going to work well on modern monitors). That'll likely be a limiting factor, especially because....

It's bad. Real bad. Like 4 hours bad. I could potentially squeeze out another hour out of this thing if I were to replace the battery (iFixit sells replacement parts for this), but it might not actually be worth the effort until the battery gets worse.

The screen is a little small for 2024 and a lot of things assume you're going to be on a larger (or much smaller) display if you're on a computer and not something like a tablet with a touch interface. The resolution is also kinda low, but I didn't find it particularly problematic for my use-case which is writing.

The brightness is also a bit dim, which would make using this outside in bright light fairly difficult, which is a bummer since it's so portable it's going to be a problem.

I actually really liked experimenting with this. At this point I've tried it in a couple of different

11ty and Ghost

Ghost/11ty

I run two separate blogs which I post to regularly (well, for this blog semi-regularly): This site (https://alextheward.com) and my TTRPG / LLC blog Cthonic Studios. AlextheWard is running 11ty hosted on Netlify while Cthonic Studios is running Ghost hosted on a VPS that I rent from one of several hosting providers.

😈 Really, I use Hetzner for a lot of things, but I've also got stuff floating around on Racknerd (an OVH reseller), Digital Ocean, AWS, Vultr, etc.

I'm constantly on the search for what the best writing experience is, and I recently wrote about my fiction writing workflow on the other blog and I want to take a bit of time to talk about the blogging workflow. Which is different! For “reasons”! It's also a good opportunity to talk about the strengths and weaknesses of both platforms especially in light of the current WordPress explosion happening.

😈 Seriously, it's bad for everyone and no one is coming away from this well off.

So, I guess this is a bit of my contribution to "what do you do if you want to produce content on the internet and WordPress is no longer a viable option, or Drupal is too complex to host and maintain"?

Let's dig in.

11ty logo

11ty is one of many static site generators (other notable SSGs: Jekyll, Hugo) who's primary goal is to take a variety of files and transform them into a website that you can then host cheaply wherever. Because the output is just HTML files, it's extremely easy to serve and there's basically no attack surface to exploit (unlike a heavier CMS like WP or Drupal).

11ty in particular stands out because of its strong commitment to stability. You can leave an 11ty site alone for a long period of time and when you come back to it, it'll still build with minimal fuss.

Don't just take my word for that though, here's a post from a friend about moving a Drupal site over to 11ty (and a Django site over to Jekyll) describing several of the same considerations: https://digitallatin.org/blog/new-site.html

This is the 2nd SSG I've used for this site, first was Metalsmith (coming from Drupal) and now 11ty.

What about it is so good, you might be asking?

Of all the SSGs I've used in my personal and professional life, 11ty has been my favorite for several years now. It's really good. I also appreciate just how little I have to worry about the site completely exploding during an update. It's a tank, and once I got the proper CI/CD workflow running, it's pretty set-and-forget.

That said, it's not very ergonomic when you compare it to something like Ghost's editing interface.

I think one of the biggest challenges for running something like an 11ty blog when compared to running Ghost, especially if you're looking to monetize content (like, running a paid newsletter), is there's a lot of work ahead of you to get it working. Even if you're not looking to monetize, but want to have on-site engagement like comments the static environment presents considerations and challenges.

Search, even, is something you have to deal with separately. You can do what I did, and use lunr.js to create a little json search index and use that. You could also use a service like Algolia, or even just Google Site Search, but there's nothing that's just there for you like you might find in a traditional CMS.

Likewise, post creation suffers a bit from just being markdown. All of your post controls are handled in the front matter of a post (I'm glossing over a lot of options you have though, check the docs).

If you want neat effects or callout boxes you're going to have to write those yourself. If you want a quick way to do embeds or include web components, once again you're on your own (and will likely need to rely on snippets).

But what about transitioning from a traditional CMS?

Permalink to “But what about transitioning from a traditional CMS?”

But what about those folks who are coming from a traditional CMS that have no idea about a development lifecycle, no clue what git is, no formal introduction to deploying sites?

It'll be an uphill battle for those folks who don't have a technical background, I think. In order to get an 11ty site up and running you need to do / learn the following:

  1. Learn how to use your OS's command line.
  2. Install NodeJS, NPM, and Git.
  3. Clone your starter repository (after learning what git clone even means)
  4. Figure out where you're going to deploy the site (11ty's docs give you some good ideas)
  5. Write your content, probably in Markdown (you do know what markdown is, right?)
  6. Run the 11ty server, behold your new blog.
  7. Figure out how to get your generated site onto your host (depends on your host).
  8. Maybe set up CI/CD to deploy stuff
  9. Maybe figure out the gaps where your selected template doesn't cover, like maybe search.

It's a lot to ask of someone who just wants to write and share their writing.

😈 The 11ty community is lovely, and I know there are folks who want to make static sites more ergonomic.

There are a couple of tools like Frontmatter CMS or Publii trying to bridge the gap a little but they both have some quirks.

Frontmatter does help the editorial experience a bit

Permalink to “Frontmatter does help the editorial experience a bit”

In fact, I'm currently trying out Frontmatter CMS for this post, and I really like several of the features. Being able to see the posts listed out in a dashboard and filter them easily is something I've been missing from Ghost over on this side of the house. Likewise, managing media manually is a pain, so I appreciate having a useful tool to help post authoring.

Frontmatter Post Dashboard
Frontmatter dashboard

But a lot of the settings are still handled in a .json file. Many of them are confusing. It doesn't help with the actual publishing of the site (which Publii tries to do, but I think could be less confusing, too). We've got a fair amount of work to do to get static hosting on par with something like a managed WordPress site.

😈 Sidebar: Frontmatter has “AI” features. Which I hate. They're disabled by default, which I appreciate. It's kinda ridiculous that it's got a “help bot” for learning Frontmatter features, just make the features easier to understand! Don't give me a hallucination engine.

Frontmatter Media
Frontmatter Media Dashboard

Ghost, by contrast, is really geared towards bloggers and newsletter authors who also want to potentially monetize their content.

As such, it's got a lot of features right out of the box to support that endeavor, and it guides you through setting them up right after you install it or sign up.

😈 Ghost.org has a paid service which includes hosting, but the OSS project can be self-hosted. I'm running my copy in a Docker container. We'll talk about the disadvantages about that.

  • Out-of-the-box configuration for the most common workflows: Blogging and sending out email newsletters.
  • Integrations! Webhooks for custom workflows (I use this for Mastodon cross-posting) and Zapier, among a few others.
  • Stripe integration for accepting payments both for subscriptions and as a tip jar (if you want).
  • Authentication for subscribers.
  • Private posts
  • The editor is "markdown plus" where you can enter markdown text and it'll render automatically, but also you can press a button or type / and get a menu of common embeds.
  • Lots of themes available for install right from the interface
    • Customizing the theme, however, is a bit of an exercise in editing and uploading code. Getting a dev environment working is also… fun.

I think the biggest thing, though, is how much configuration you can do without having to touch the code. Need to add analytics (like Plausible)? There's a section to add arbitrary code to the header or footer.

Generating menus is easy. Most themes have slots for where those menus go.

Search is included (though it doesn't index the entire text of a page), no config needed.

Scheduling posts just works! No need to set up a CI/CD workflow or a cron job to look for and publish new posts. It's all included.

😈 Though it's important to note that is a cron that Ghost provides...which hits the external URL of your site. If you've got, say, bot protection turned on, your cron might get blocked and not make posts live.

But I think the best part is the editorial experience, which I've been chasing on the 11ty side for a while now.

I want to take a bit of time to highlight the thing I like the most about Ghost: The Editorial experience. For me, it strikes a very good balance between raw markdown and the convenience of a CMS (because it is a CMS).

Ghost Editor
Ghost Editor with sidebar expanded

This winds up being similar to how Obsidian feels when editing, but the additional embeds and "extras" you get make it feel smooth. For example, I do a fair number of product reviews / game introductions over on that blog, and having a "product card" convenience element is very useful. The callout element is simple, but I make heavy use of it on both blogs.

Product Preview Element
Product preview element

Like Frontmatter, it has a series of boxes for the SEO settings per post, as well as share previews for when you're publishing your post. It's tight, and has the features you'd expect from a CMS.

Alright, where does it fall behind SSGs like 11ty?

Permalink to “Alright, where does it fall behind SSGs like 11ty?”

Simply put, maintenance becomes your problem if you're not using Ghost.org's hosting. Relatedly, configuring the site to handle traffic also becomes your problem. For Cthonic Studios I had to configure the following (via code):

  • Backups of the database and file system from the docker container.
  • Uploading those backups offsite from the VPS (in case the VPS explodes).
  • The CDN to keep traffic from overwhelming the origin server.
    • Likewise, ensuring logged-in users do not get served the same cache.
  • Periodic updates (which is usually just a backup / docker pull)
  • Monitoring (it can go down! That'd be bad)

With the 11ty site, I don't have to worry about backups at all. It's all markdown files in a distributed git repo. If the server goes poof, everything's safe and sound in several places (that are themselves backed up). I could have the 11ty site back up and running in a few minutes if I needed to switch hosts. Ghost is probably going to take 30 minutes to an hour to restore from backup.

And for someone who's coming from another CMS?

Permalink to “And for someone who's coming from another CMS?”

I think it's probably obvious here that if you're coming to Ghost from another CMS (like WordPress) it's going to be an easier and more familiar editorial experience than trying to set up a static site on your own. As nice as 11ty and its peers are (and rock solid), the end-to-end experience of getting running on it is filled with a lot of things that require more tech setup to get running.

😈 This isn't to say that someone couldn't make a rock solid static CMS setup experience, and Publii seems to be trying, but we're not there yet.

Ghost, for its part, is more akin to something like Buttondown than it is to Drupal or WordPress but I think a lot of content-first sites would be well served by moving over to Ghost.

I think the main issue here that all the WordPress stuff is dragging up is both ownership of your content and portability of that content. A major advantage of the static site path is portability. All of your content is just in text files, that can be processed in any manner, and then placed somewhere else.

This is one of the things I'd like to see Ghost do better. You can export all of your posts in a giant JSON file, but doing anything with that is going to require some development work. Likewise, pulling any images down will require some effort to get the content back out. There are some tools that can help that conversion out there, however.

I think I'm reasonably content with my current blogging setup aside from trying to get a better editorial experience over on the 11ty side of the house. This afternoon I'm planning on actually setting up a Gitlab CI/CD workflow to rebuild the site so that I can pre-publish posts like I can on Ghost (yeah, I never did set that up) but I also want to have the ability to publish down to the hour...and I don't want to do a deploy unless something's changed. I suspect I'm going to need to do some wrangling to ensure that works like I want it to.

Either way, I like both 11ty and Ghost and I am using them for different enough things that I think I'll just keep going as-is. If I ever decide to try to do more subscriber-heavy content over on the AlextheWard side of the house I might migrate it to Ghost or I might just use that as an excuse to talk about using edge functions for dynamic content.

We'll see. Hope this was helpful. Happy blogging, friends.

Riding Amtrak from PDX to LAX

Superliner Train Cars

I've always wanted to travel the slow way across the country, taking a train somewhere overnight. Watching the scenery go by the window when I'm in no particular hurry to get anywhere. I had the opportunity to do just that this past week riding on the Amtrak Coast Starlight down to LA to visit the in-laws. Obstensibly, I can call this a research trip, as my in-progress TTRPG Nix Noctis (working title maybe permanent title) features a post-post apocalyptic train line that encompasses the southern half of the route. What better way to ensure I've crafted an "accurate" half-passenger half-freight half-military rail line.

Regardless of justification, it was quite the experience. I booked a roomette with Amtrak months in advance (probably too far in advance as I don't think I got the best rate. Oh well, experiences are worth it). Since my wife wasn't making the trip with me, I had plenty of space to stretch out in the sleeper, though I do wonder if there'd be a good enough space for the both of us in here for 30+ hours. I suppose there's always the observation car, but it was full basically every time I visited it.

Now, I've ridden Amtrak a bunch of times in the past, but this is my first time on an overnight journey. The longest trip I'd ever taken was to Vancouver BC, at just over 10 hours due to multiple delays. Thus, those trips were always in coach. This time, since I was going to be on this thing for over 30 hours and ideally I'd like some sleep that isn't surrounded by other people, I figured this was the best time to do a sleeper car journey.

Clouds and fields

Right when I boarded the train I was greeted by our car's attendant. He informed me that I should promptly drop my stuff off in my roomette and head to the dining car so I could grab some lunch before they closed it down. I took his advice and headed that way. Since I was traveling alone I was paired with another solo traveler for the meal, which was something I was mentally prepared for but not exactly expecting initially. I decided to have a Beyond burger since I was kind of hungry, but knew dinner would be coming up pretty soon. I managed to finish my meal just as we were started to pull out of the station. Eager to start getting some writing done, I headed back to my roomette and did a little bit of setup.

My home for the next 30 hours

But I totally forgot my burger came with a brownie, so I missed dessert! The dining car attendant offered to make me one but I politely declined as I knew dinner would be sooner rather than later.

Back in the room I set up the little fold out table into a small workstation. It's basically the perfect size and distance away from me for a 13" laptop, but I think a bigger laptop would require some adjustment or simply using it from your lap. I watched the familiar sites of 99E roll by as we followed along the highway I prefer to take when heading down towards Salem. It's both weird and comforting to see those sights when I'm not in control of the moving conveyance. Relaxing almost, but with a sense of being out of place.

The room itself is fairly small, but comfortable for me (I also really like micro hotel rooms where there's not a lot of space). The main downside to the train car is there's just a single 120V outlet, which means I had to be a little creative. Thankfully, I have a high capacity Anker battery pack which I can use to charge both the phone and the laptop for a bit and then charge the battery pack from the outlet when I'm out of the room. The other bit is there's no way to lock the room from the outside so really you just want to keep your belongings either well hidden or with you.

That said, when you're not at a stop there's not many places you can go and there are only so many people who are going to be in your sleeper car.

I sat and enjoyed the views while typing out several posts and story revisions that I'd set aside from the trip, before dinner my word count was sitting at right around 3000 words, not bad for about 3.5 hours of writing and exploring a train. As a writing nook, I give the roomette a solid A+. It was very comfortable on my back and wrists, and I felt very relaxed and productive. As the sun started to lower, I had to switch to facing the other direction so that I wouldn't have it directly in my eyes, but a quick move from one side to the other was all that I needed to keep on writing.

As we left Eugene and headed into the moiuntains, the scenery became quite lovely and the cell signal dropped well out of view, just in time for me to step back into the dining car and make some new friends.

Like lunchtime, I was seated with some strangers: a couple and an older gentleman. It was a fine time, and I ordered the steak. Unfortunatley I forgot to snap a photo of that before I started sawing into it (I was hungry, okay?). Each dinner meal comes with a complimentary alcoholic beverage, so I paired the steak with a glass of cabernet sauvignon. The steak was impressive, well cooked and juicy. Paired with mashed potatoes and green beans. It was tasty and filling. For dessert I opted for the cheesecake, which was also quite good.

Movie time

After dinner, I grabbed another (paid) beverage from the cafe car, located just past the dining car under the observation deck. Then I headed back to my room to watch a movie and play a few games on my Aya Neo Air before bed. Once I was ready to sleep, I pulled on the little "call attendant" button in my room to have our cabin steward friend come set up the bed. The whole process took just a few minutes and then we were in "sleep mode".

The bed setup and linens were quite comfortable, though in a roomette the amount of space you have to maneuver when the bed is made is... limited. Basically just a small square at the front of the room. But I did have enough space to pull the tray out and set my iPad Mini on it to keep watching my movie. That was cool. After the movie was done, it was time for some light reading and sleep.

Superliner bed all made up

I did not sleep well at all on this ride. There were a lot of issues with the sleep process. The first: the hallway lights do not turn off and the curtains do not block all of the light. No problem, I packed my manta sleep mask (which I love, it's amazing) so if I couldn't block the light out at the source I'd just block it out at my eyeballs. The second problem was the sound of the announcements. The room does have a knob to turn down the volume that I used, but near as I can tell the "ding" of the call attendant button isn't covered by that setting and several someones pulled that knob throughout the night. Finally, temperature control: the comforter that was provided was amazing. Felt wonderful, but I was unable to properly thermoregulate with it on so I wound up pulling it on and off throughut the night, and this is with my room set to the lowest temperature setting on the dial.


Curtains, but not blackout curtains

Either way, I did manage to get some sleep; I didn't particularly mind the rocking back and forth though there were some times where it felt like the train was going full tilt and hitting some bumps along the way. Every creak and groan you hear during the daytime is more noticeable when you're half-asleep.. We stopped for a long period in Redding around 6:00 AM, but I went back to sleep until my regular alarm at 7:30.

After I got up and smelled coffee brewing in the hallway (or more accurately the coffee had been brewed in the dining car and brought to our sleeper) I got dressed and went to check out the shower area. Blessedly, the shower was presently unnocupied and hadn't been used today so I took full advantage of it. They provided soap, conditioner, and body wash as well as towels. The shower was great, had hot water and everything. Just what I needed to refersh after a night of poor sleep.

After that I made sure that my door was left ajar for the cabin attendant to see so he could make the room back up into chairs and headed to breakfast. Like every meal before it, I was seated with some nice folks this time who didn't know each other. We exchanged some plesantries and enjoyed some amazing views off of the right side of the train.

It's worth lamenting at this point that my side of the train faces inland, so I do not get to enjoy the wonderful costal views from my train car. I tried to get a spot in the observation lounge, but alas it was full.

For the meal, I opted for the French Toast. This turned out way better than I was expecting, it was delicious. The fruit on it was particularly fresh. I don't know if I was just super hungry or what, but it tasted great.

Breakfast french toast

After breakfast, I headed back to my room to get started on some more writing and to recharge my battery pack. This is when I learned the horrid truth: my charger brick wouldn't stay plugged into the provided outlet. Since I'd charged my phone overnight the battery pack was sitting just over 55% charged and presently have no way to recharge it. Fortunately, I didn't need the power because I'd brought the Macbook air with me and after a day of heavy usage it was only sitting at 65% battery. I still had plenty of battery left to do the writing I wanted to get done and probably watch another movie on it if I so desired. Even if not, the battery pack was charged sufficiently to top it up. Since I was only on the train for another 11 hours at this point, I figured I was in pretty okay shape.

Plus I carry a backup magsafe phone battery for phone emergencies.

Otherwise it was a pretty uneventful journey working on some more blog posts and thinking really hard about how to and if I even should revise my second fantasy short story (an additional 2000 words written from 9) when I break for lunch back in the now-familiar dining car.

Though apparently they oversold coach and there weren't enough seats. Whoops?

Just before I was called in for my lunch reservation time, an announcement came over the intercom. Apparently, just ahead of us there had been a high tide and we had to wait for the water to recede and a quick track inspection to ensure nothing was going to go kerploof when we went over it. As such, the view from my dining experience was, well, this:

The view out my window.  Brown dirt on a mostly empty lot

Regardless of the view (which did improve shortly after we got rolling again), this time I went for the patty melt which was just as brown as my surroundings.

Lunch, a ruben sandwich

But, like every other meal the overall quality of the food was solid. The bread was buttery and crisp and the meat was tasty. For dessert I had a bit of a chocolate brownie (the dessert I'd neglected to eat the day prior) and was duly impressed with it. After finishing up my meal and excusing myself from the table I headed back to my cabin where I plopped down and immediately went into "digestion mode". Which is to say, my brain shut down and I stared out the window for a good hour while I regained my bearings. The lack of sleep from the night before was starting to catch up to me.

The scenery rolling by outside is a stark contrast from what it had been at the beginning of the trip. Where a good portion of Oregon was covered in greenery, California was covered in rolling hills of short shrub grass and sandy dirt. Interspersed between those are vast farms growing a bunch of stuff. I think we passed a wine grape farm at one point. Makes sense for the part of the state we're travleing through.

Fun fact, apparently Soledad, CA has a very chill set of weather patterns, with the average temperatures in the 60s and 70s, always.

For writing, this block of time was a bit less productive than the morning due to the fatigue and digestion catching up with me, but I did manage to churn out a solid portion of a blog post while also fixing the SSL cert on one of my servers. Sidebar: the cellular connection has been pretty consistent this whole time, keeping me connected now that Amtrak no longer offers wifi.

Probably the coolest part of traveling south so far has been the very hilly region that the Amtrak goes through on its way to San Luis Obispo. It's just this winding path on the upper portion of one of the hills, a trip through 4 or 5 tunnels until eventually you wind up in the town (the last fresh air stop before LA).

Valleys and Hills

Part Five: Dinner, A Situation, and Arrival

Permalink to “Part Five: Dinner, A Situation, and Arrival”

I chose a relatively early dinner reservation for this evening, mostly because I wanted a lot of time to digest before we arrived in LA (and since I was getting off a few stops early that seemed prudent). I forgot to get a picture of this meal (whoops) but this time around I had the salmon entree. It was great. Not a giant portion, decent serving of vegetables and rice with it, and a nice semi-spicy lobster sauce. It was great. We hit the ocean area just as I was wrapping up my meal too, which had some stunning views of the Pacific ocean (which I didn't get to photograph as I was at the table with complete strangers and thought that would be rude).

Unfortunately, I skipped dessert because there was a medical emergency on the train, and we had to make an emergency stop in Lompoc-Surf where paramedics were waiting. I didn't hear the full story, but I hope they're okay. The unexpected stop did afford me a stunning view of the sunset.

Stunning view of the sunset over the pacific ocean

For the rest of the evening I plotted out a little more writing, and settled on what I wanted to rewrite on my story (but I hadn't actually started doing the writing just yet). But as night fell I just wanted to turn my brain off a bit, as it'd been a long day on very little sleep.

We pulled into Van Nuys and I hopped off, saying farewell to my "home" for the past 31ish hours. It was a great trip and I'm glad I did it.

There was a lot to like about this trip. I was able to get a ton of writing done and feel fairly proud of myself for remaining as focused as I was (even though I didn't do as much as I would have liked. I'm never satisfied with my own output). The food was really good for what I was expecting. It's not a five-star restaurant but it is very solid for what it was. I wasn't disappointed in anything I ate on this trip: especially since it was a solid five meals.

The views outside were absolutely gorgeous, and while it'd have been nice to be on the ocean side I did still get some really great inland views instead. In fact, it was really great to do this route and I have some more descriptive stuff ready to go for the portion of the game where this rail line is concerned.

Really the only downside I can think of is how difficult it was for me to get to sleep and the resulting tiredness on day two of the trip. And maybe the outlet. I might bring a charger that has the ground plug next time so it'd stay in. But that's about it.

Also money-wise this wasn't too terrible. I wound up paying $620 for the roomette from PDX to LAX (though I did hop off a little early). That included 5 meals that would have otherwise been charged, totalling $165. Definitely more than you'd pay for a flight, but still under what a first-class ticket is going to run on most domestic days. Plus I got a ton of writing done in the time, and that was well worth taking the extra time.

Ergonomic Static Site Blogging


Shutterstock #2298461049 by One Photo

I've been blogging a lot the last couple of weeks across both 11ty and Ghost, and it's no secret that I really like Ghost's authoring experience. I'd love to replicate that in a static site, and get a little closer to functionality parity, but I'm still searching for a great way to do this.

😈 Sidebar, I love everything about that stock photo. Not only is it a tablet with an Apple Keyboard in front of an apple laptop, there's a calculator to the side, a notebook, and what appears to be a plant potted in a cup? With CONTENT just screaming at you. The composition is wild.

This post is going to be some in-public documentation about what I'm trying, and how it's going. For the "why" portion of that question, hop on over to my previous post about 11ty and Ghost.

The search for an ergonomic blogging experience

Permalink to “The search for an ergonomic blogging experience”

When I started this post I had the following ideas in mind:

  • Just using Ghost's API and hooking it into 11ty
  • Checking out Publii
  • More tinkering with Frontmatter CMS
  • Some other, more paid-er, option?

So let's hit those points one-by-one.

This option is a very interesting one which Ghost has written about already, and it's an intriguing one. It basically sits in the same space as using other CMSes headlessly. You'd:

  1. Stand up a Ghost instance
  2. Expose the content API
  3. Have your 11ty build step pull from the content API to generate the static pages
  4. Deploy however you'd like to

This all works just fine, and with the Content API you get some features of Ghost (drafts and future-dated publications) without having to wire up the collection filters in your .eleventy.js file. But that's right about where the benefits end.

For one, you now have a fully functional Ghost site which you could be serving traffic on behind a CDN for a similar experience (which I wrote about on Cthonic Studios). Sure, the static pages are always going to scale better than a full on server, but you can get to "pretty good" quickly with Ghost's headed mode.

For two, to do previews of the actual site you're now beholden to a two-step process where you draft the post on the Ghost side and then run an 11ty server to see the preview of it when it's actually rendered.

You also lose the authentication features that are built in to Ghost, having to reimplement them at the edge. In short, this has all the same downsides of a headless CMS (like Strapi or the like). For me, that's just not worth it. I'd rather start from something that has markdown and can be rendered by itself in a nicer editor.

On towards the next idea:

All of this is on Publii v 0.46.1

There's a lot to like about Publii, it's a desktop app that puts the rendering side of the site on your local machine that then pushes to the server of your choosing. It's got a lot of GUI-configured deployment targets, like FTP, GitHub Pages (soon to just be "git repos"), Netlify, S3, and so on.

Under the hood it conceptually works like most static site generators. There's an input directory which it uses to generate the flat files, moves those to an output folder, and then syncs that folder elsewhere. The actual structure of the site is determined by the Theme, which you're free to create yourself and attach to the site in question (like Ghost).

Publii Dashboard

Though, it does store your post content in a SQLite database. Which is... neat?

Publii SQLite DB

I'm actually a little intrigued by this data structure because it changes depending on how you're authoring the post. If it's being done in Markdown, then the markdown is just there in the post. If you're using the block editor, JSON. This theoretically means you could reproduce the Markdown or do other shenanigans in a plugin or somesuch, I want to dig into plugin authoring a bit more and see what can be done with this.

But, back on track for the moment, the idea for less technically inclined users is you can just fire up the app and get to blogging. This is relatively true and for me a good experience. It offers 3 different editing modes for posts and pages:

  1. Markdown
  2. Block
  3. WYSIWYG

Publii Block Selector

Each of those works well, but if you just use the Markdown editor, you're missing out on conveniences that you'd find in other editors (like the ability to use the image carousel). They also have variant features. For example, the WYSIWYG editor does word and character counts. Neither the block nor Markdown editors support this.

My biggest annoyance about the editors is the fact that it does not autosave nor can you press a hotkey to save the post. You can only save your work by clicking on the save button. That's wild. I suspect the hotkey would be easy to add, but there must be some nuance there otherwise they'd have done this already.

Menus are straightforward and there is good integration with the theme settings. When you create a menu, the locations of menus from the theme are populated in the dropdown:

Menu Selection

There's a lot of other low-code stuff buried in here, colors, theme settings, fonts, etc. A lot of setup can be done without ever touching the theme files or adding custom CSS (though of course you can do all of that, just like in 11ty or Ghost).

Overall, the experience shows a lot of promise, but has a lot of rough edges that make it harder to use. For one, if you're editing a post, and you preview it, it only generates the file you're on. You cannot navigate around the site. To navigate the whole site (for example, to check the listing pages) you must go back to the dashboard and click the preview button there. It's annoying to break the flow like that.

A major omission that I'm not sure how I'd accomplish is the lack of scheduled posts. You can set a post date in the future, but it will appear on the site immediately when you sync it. There's no cron or anything that would allow you to automatically show a post when its date has passed. This makes sense, because the editing and publishing happens from the app, you'd need to keep it open on a computer that is not sleeping in order to do that. It doesn't look like this is a thing the developers are interested in supporting either, given the complexities.

Implementing something like that would be tricky, but potentially doable via a plugin and a particular provider, but I think it'd be a fair amount of work.

The other thing is the plugin ecosystem. It's pretty robust and supports paid plugins, but there's nothing in there that I can find which supports members-only pages or the like. That might exist in the future, but just like other static sites it would depend on edge functions and a provider that supports it. So it's unlikely it'd make it in.

As a replacement for my 11ty site, the lack of a future-publish feature and the lack of control makes it unlikely that it'll work for me, but I may use it for other one-off sites, it's pretty neat.

So, this is what I'm using to work on this post right now. It's a VS Code Extension that adds some nice convenience features like Easier media management and a nice sidebar for managing the frontmatter of posts. It supports 11ty pretty much right out of the box, and has buttons that hook right into 11ty's lifecycle.

Frontmatter CMS Dashboard

So far, I think this is going to be the best way to accomplish some of the things I want to do, but that frustrates me a bit since I want to reduce my dependence on VS Code, not least of which because I expect Microsoft to add more AI features (which can be circumvented via VS Codium), but more on that in a sec. It's also not a great solution for someone who's just looking for an enjoyable static CMS.

That said, let's talk features.

After doing an initial setup and telling Frontmatter which folders your posts are in, you'll get a clean dashboard which lists off every post you've got, along with their status (draft, published), cover images, preview text, and the like. Right next to that is a Media Editor which is really just a file browser with some convenience and insert functions.

Frontmatter Media Page

Next to that are snippets, which are just little hunks of code you can use to easily insert things.

😈 For example, I made a callout block snippet which made this block. No more quotes for me!

It does insert a fair amount of comment blocks for...some reason? It looks like they're present so you can then edit them in the snippet inserter again later, but I can't figure out how to do that.

Then, there's a whole "Data" section which allows you to do things with a data cascade, and I think you can hook that directly into 11ty's data cascade for fun and profit.

Finally, there's a nice little taxonomy editor which is awesome because it automatically parses the tags you've used on other posts, and correlates them to your posts. It's just really handy to get an overview of what tags you're using where, which is hard to do with just vanilla flat files.

This ticks a lot of boxes for me, and it'll likely be what I use until I find or build something better, but there are a couple of problems that make this hard to use for a less-technical user.

  1. A lot of settings can only be edited in a json file. There are a lot of UI controls missing for changing some settings you will definitely want to change.
  2. Editing media in the media dashboard does not edit media in the post, you have to reinsert it.
  3. Create Content settings default to prefixing your file with a date/timestamp, which is far from my preference (but you can change it via the aforementioned JSON).

All of these are relatively minor, but there's one thing that annoys the heck out of me.

It has "AI Integrations". Now, these are disabled by default (good), and premium features (ugh, why would I pay for that garbage?) and they seem like they're just bolted on to make a specific crowd happy. They also include a button to "search the docs with AI" called "Ask the Frontmatter AI for help". This robot comes with a disclaimer:

Warning: Answers might be wrong. In case of doubt, please consult the docs. - Frontmatter AI

Note, that "Answers" is not my typo. It's in the AI response. This feature is also absolutely worthless. Here's a screenshot of me asking it about Tuna and it responding in Spanish with a warning about its inaccuracy.

This is a useless bot

Why the absolute hell does this exist? You could have just done a search for me and presented the results. You could even do a vector database so you can use better natural language search. Instead, you get an LLM to spit out a poorly summarized response, possibly filled with inaccuracies, that costs you money, me time, and the environment 4 bottles of water.

😈 Seriously, I feel icky that I even used that thing. I'm going to need to plant a tree in penance.

The fact that something like that is included just strikes me as really poor planning, or an ill-advised play to get AI bros to give you money. I'm not sure which, but it makes me irrationally angry. We really have to stop doing shit like this.

So the only other things in this space seem to be stuff like Siteleaf (which is a thin wrapper on top of Jekyll) or Drupal + Tome + some more setup, or Strapi + some other frontend. I've done a bit of research in this space but haven't found anything really compelling.

Tell me your favorites! Yell them at me on Mastodon @cthos or via the Contact Form. Maybe one day I'll set up comments over here, but that also involves comment moderation.... but also easier to hear from y'all.

But yeah, that's where we're at. I'm sure I'll have more of this as I go. I'm slowly thinking if I want something pretty perfect I'm going to have to build it myself, and I'm not sure if I want that.

Auth Series 1: Sessions and Passwords

Auth Post part 1: Sessions and Passwords

Hello and welcome to the first part of a multipart post series where I'll dig through the depths of authentication knowledge and share all of that with you, including some deep lore of how we got from the early days of the internet to today (though, I'm a millennial, so I wasn't exactly a web professional back then. Neopets all the way).

😈 This guide is targeted primarily at web developers who want a deeper look at authentication implementations and some (but not all) of the considerations that go into them.

Part 1 will cover the following:

  1. How do web servers identify requests?
  2. Session storage and state management
  3. Passwords and how to store them as securely as possible
  4. Handling forgotten passwords
  5. Signing out and Account removal considerations

I've also created a repository on Gitlab with a couple of examples for session state: https://gitlab.com/cthonic-studios/authentication-examples

Now, there are going to be a lot of caveats to this topic - authentication is a deep topic and has a lot of gotchas. I'll cover some of them, but I'd need a full book to cover everything. That said, let's get right in!

😈 This post covers traditional multi-page apps being served by a webserver. We'll talk a lot more about Single Page Apps and their variants in a later post.

This post assumes some familiarity with web servers and web technologies. HTTP, cookies, methods of data storage, the concept of multiple web servers working in coordination. If you need some grounding in web servers, Digital Ocean has a good guide.

Web Servers are stateless. How do they identify requests? Sessions!

Permalink to “Web Servers are stateless. How do they identify requests? Sessions!”

So the very first thing you need to consider when dealing with web authentication is that web servers are stateless (more specifically, the HTTP protocol is stateless). Whenever your browser makes a request to a website, the server has no idea that one request is connected to another. There will be some telemetry, like your IP address or the browser you're using, but as far as the web server is concerned, each request is totally independent of any other request.

This is a problem if you want the server to understand a given series of requests are actually related to a given account. Without some unique identifier to pass along to the server, you can't associate a given page load with a given account, which makes persistent web applications kinda difficult.

To deal with this, we have the concept of a browser session. The very basic premise is that the browser will store some bit of information and send it along with every HTTP request that the server can then use to uniquely identify the "session" in question. There are several ways that one could do this so we're going to talk a bit about some of those techniques first and then talk about best practices.

I'm going to use PHP's session management as an example because it has a lot of built-in functions for managing session state, and they're quite illustrative of how it works under the hood. They even have a best practices page for proper session storage.

So, to start, PHP has a function called session_start() that will initiate a session for a given request, which allows you to store parameters in the $_SESSION global which will persist across requests. Just like magic, you've got a storage area where you can store information about the session in question. But how does this magic function work?

The very first thing PHP does when starting a session is generate a random session_id() for the session, and sends that back down to the browser in the response. How it sends it depends on the configuration. By default, it sets a cookie via the Set-Cookie response header, with the name of PHPSESSID. There are a number of configuration values for this, allowing you to disable session cookies, control their behavior, change the name of the session variable, and so on.

The generated session ID was a derived ID based on a number of factors until PHP 7.1, but now it's generated using random_bytes. Either way it results in an alphanumeric ID which isn't guessable in a reasonable amount of time given current computing power.

We'll talk about the cookie thing in a minute, but for the moment let's talk about what happens if you don't send a cookie back. How else might we get that session information? Well, if you disable the cookie setting, and enable session.use_trans_sid, the session information will be transparently appended to every anchor tag on the page so that ?PHPSESSID= will be appended to the target URL, thus preserving your session across requests through the power of $_GET variables. If you'd like to see this in action, I've included some simple examples in the aforementioned code repo.

Now, this isn't great, because that makes your session ID directly visible in the URL. GET params are encrypted over https:// but anyone who wants to look over your shoulder could take a picture of your session ID and then hijack your session simply by using the same ID. This is why basically all modern systems use Cookies to store your session identifier. We'll talk about GET parameters more when we discuss other auth systems (OIDC and SAML both use GET parameters for various things), but for now, just internalize "use Cookies to store session info".

😈 Why didn't we just always use Cookies? Why does this setting exist? Well, in the olden days, Cookies weren't always reliably available. We live in better times now, kinda. Wait until we talk about 3rd party cookies.

Okay! Now that we know how to start a session and how the server identifies a stateless request, what can we do with that? Well, now you can store information about that user in the session storage. In PHP, you can directly assign things to $_SESSION['example_key'] which will persist across requests. PHP's basic session storage is file-based (meaning it stores your session data on the hard drive of the server you're interacting with), which is generally unsuitable for distributed applications because you'll have to either ensure the user is pinned to the server where their session information is stored, or* you'll need to synchronize the file system across all web servers. Neither of those is ideal, but it's definitely a thing we used to do many years ago.

So, you can also store your session information in some other system, like Memcached, or a database which is what modern systems tend to do. Indeed, if you're working in a language without a built-in session mechanism, this is what you're going to want to do. Let's talk about that a bit more.

Storing session state is a complicated topic and depends a lot on your server topology. If you just have a single web server, storing state on the file system is going to look "fine". If you have multiple web servers handling requests, you'll need your state to be available to them. There are several ways to handle this, and if you want a quick overview of the options Laravel's session storage page has a good series of the typical options. Including:

  • Storing data on the file system on a single server
  • Storing data on a networked file system (block storage on a cloud provider, typically)
  • A database, either the application database or a separate database
  • In-memory storage like Redis or Memcached
  • Directly in browser cookies, typically signed or encrypted.

The most common options I've seen implemented for session storage are either "just store it in the database" attached to the session identifier, or "store it in memcached/redis" also attached to a session identifier. Over the course of my career I've seen the progression from file → database → memory cache storage several times now, and that's directly correlated to scaling applications. File storage (even on a networked drive) is slower than the db which is slower than memory, so when you're wanting to squeeze more performance out, this is where you tend to go.

The other option which requires very little configuration is an encrypted session cookie. Most frameworks handle this for you because if you put session information in a cookie the client can tamper with it, which is usually something you don't want to have to worry about.

😈 You could also use file-based session storage and just ensure that a given session is always served by a given server with some load balancer trickery... but that's a topic for another day.

That said, you also generally want to minimize the amount of data you're storing in the session in the first place.

At minimum, for an authenticated user, you'll need to store their user_id or another primary key so you can identify that session to the user in question. We'll talk about that more in a second.

For anonymous sessions, you might do something like store a shopping cart (for e-commerce) or preferences, or any number of other things that you need persistent across page loads.

This can also be things like error or status messages that need to be shown on subsequent page loads, for example if there's an error on a POST request and you need to redirect back to the prior page with a message. You could include that in the URL, but it's commonly stored to session storage instead (because someone else could send you a link with a misleading message via that same URL).

This really depends on your application, but I'd recommend storing as little data in the session as you can get away with.

As the OWASP Session sheet calls out, you have a few different options:

  • Keep the session alive indefinitely, until the user takes an affirmative action (logging out)
  • Expire the session after a period of time.
  • Expire the session after the user has been inactive for a period of time.
  • Expire the session when the browser is closed.

Which one you'll choose depends on your application and how sensitive it is. I've encountered a lot of blanket "you must log the user out after 15 minute" corporate requirements, and I need to tell you that is less secure than you might think it is, especially if it isn't an idle timeout.

In general, I recommend:

  • Expire the session when the browser tab is closed (expiration time of 0 on the session cookie), but give the user the option to stay logged in across visits. In that case, set the expiration of the cookie to some sensible time in the future (weeks / months).

For more sensitive contexts:

  • Expire the session after a sensible idle timeout (yes, 15 minutes works)
  • Do some basic anomaly detection (IP addresses changing suddenly, etc), require re-authentication for sensitive operations, etc.

What about client side LocalStorage or IndexedDB or sessionStorage?

Permalink to “What about client side LocalStorage or IndexedDB or sessionStorage?”

A popular choice for storing ephemeral data on Single Page Apps (SPAs), which can be a good option for data that's tolerant to being modified by the client (meaning, the server cannot inherently trust that data, it has to be verified).

While I don't recommend doing this (use a library if you can), the basics of how sessions work boil down to this:

  1. Create a session identifier which is unique and impractical to guess (ideally long and random) or cryptographically signed.
  2. Sending that session identifier securely to the browser (over TLS, encrypted, etc.) so that it cannot be intercepted in transit.
  3. Inducing the browser to send back that session ID to the server (also securely so it's not intercepted over the wire) ideally via cookies.
  4. Storing information associated with that session in some sort of storage that's accessible by all the web servers that could potentially serve the request.
  5. Deciding how long you want that session to last and removing the session when it expires.

There's a lot of nuance that I'm eliding over in those posts, but those are the core elements of a session storage mechanism. If there's enough interest we can do a roll-your-own example of session storage, but I'd never encourage you to do this yourself other than to understand the mechanics. Much like cryptography, there are a lot of footguns, and there are many good libraries that handle sessions.

That said, for more detailed information, check out the OWASP Session Cheat Sheet.

For an even deeper dive into proper session storage, NIST's SP-800-63B document is an extremely long technical document that covers a lot of information about authentication.

You should generally have the web server manage the session with a cookie, and unless you have a very good reason to allow JavaScript to read that cookie, you should send it with the HttpOnly flag (which prevents JS from accessing it).

Likewise, you should set SecureOnly so it will only be sent over HTTPS and the SameSite attribute so it will not be sent along to other domains on the same root domain (unless, of course, you want to do that and you know the trade-offs).

Session hijacking is a technique whereby an attacker gains access to the session for a given target user and then uses that to act on their behalf. OWASP covers that in this article, but the basics are "Gain token, use token". There's also a reverse attack where you use an XSS vulnerability to inject your own session token into another user's session and get them to do some target action.

Either way, you want to prevent that session token from leaking. Here are some steps you should take to prevent session hijacking:

  • Never send the session identifier over http://, only over https:// connections. If using cookies, ensure SecureOnly is set.
  • If you can help it, do not let JavaScript access the session identifier, like with HttpOnly cookies.
  • Ensure that your session ID is impractical to guess (length, randomness).
  • Change the default session identifier variable for your language (for PHP, don't use PHPSESSID). This'll make it harder for a malicious user to guess or target in bulk.
  • Implement Content Security Policies that prevent cross-site requests to prevent JS from sending session information.

Identifying the user with Username and Password

Permalink to “Identifying the user with Username and Password”


Song_about_summer @ Shutterstock #1794130912

Right, so we've now talked at length about how sessions work, but currently we still don't know anything about the person on the other end of the internet other than they initiated a session in a browser. We can reliably identify the browser that started the session is still the browser interacting with that session, But what happens if the human on the other side of the screen needs to change computers? Or they shut down their browser? Or they clear their cookies? How do we go past identifying the browser and identify the user?

Well, that's why usernames exist. Very basically, as a website operator, we're asking you directly to tell us who you are. So, when you provide registration, you'd ask a user who they are, and that most often is a memorable username chosen by the human who's signing up for the service.

But, if all we ask for is a username, then anyone with that username is able to log in as the person in question. Because usernames are typically also displayed publically, this would be very bad.

This is where passwords come in. Passwords are an ancient invention, you might have to utter a password to a guard at a door in ancient Rome to gain access to a place. Modern passwords are directly derived from that idea, you need to give a password that ideally only you know so that we, the service operator, know that you are who you claim to be. Or at least, that you have the requisite knowledge (which is where impersonation and multifactor authentication come in).

Now you want to collect an email address and use it for account recovery or for the username. How do you ensure that this email address is an email address? You can use a regular expression to validate that it's an email, right?

No! Email addresses are...remarkably complicated in how they can be formatted. There is a regular expression you can use and catch most of the cases you'd find in RFC 5322, but you're not going to catch everything. My advice is to use a simple regular expression to ensure the format is in the ballpark (ensuring there's a local and remote part with an @, for example) and then send an email to the provided email with a one-time code in it to verify delivery of the email address. There are also services like Neverbounce that do validation by sending test emails and maintaining giant lists, but you can get by with simply making the user validate their email address.

😈 Like names, you should accept that what the user is giving you is their email address and don't try to be clever. The best validation for emails is by sending an email.

You should also do this before allowing the user to change their email, and you should put that operation behind a second factor. If you haven't implemented MFA or the user hasn't turned it on, send a code to the existing email first (following the guidelines for password resets, presented later) and then again to the new email address to confirm the change.

So, the combination of username (or email address, commonly) and password are what basic applications use to identify users with the service. Great! We can just collect that and store it directly in the database, right?

Wrong! You may be tempted to store passwords in plain text in your database, how else will you send the password to the user when they forget it? This is a really bad idea. For one, if your database is compromised, you've now given malicious users every single user's password. Humans tend to reuse passwords across services (even though we've been trying to get folks to not do that for the entire history of application passwords), which means you might have compromised a bunch of user's security across any number of sites.

We don't want that. Okay, so maybe we'll encrypt those passwords in the database! Then, if you steal the database, you can't decrypt the passwords! This is also a bad idea, because in order to encrypt the password in the first place, your application server needs to be able to access the encryption key. Which means if your application is compromised, then you can still steal plain text passwords. It might be slightly harder, but it's still possible.

So, what do we do? You hash the passwords, with a one-way algorithm so that the server cannot get the plain text password back out of it. You'll be able to check that what the user has given you matches what you have on file when you run it through the hashing algorithm, but anyone with the database cannot just get that password back out.

In the early days, we'd often just use an algorithm like MD5 without any other mechanisms in place and call it good. As computing power has increased drastically, MD5 hashes can be checked in a short amount of time rendering them insufficient for modern security. Likewise, there are entire databases of hashes called "Rainbow tables" which correlate MD5 hashes to plaintext equivalents.

So, instead, we now have better algorithms and other techniques to slow down the checking of a hash. The most common mechanism I've seen in the wild is still bcrypt, while NIST recommends using PBKDF2 which is specifically for passwords. OWASP recommends using Argon2id. These algorithms increase security in several ways, the first is they each run through a number of iterations (which is configurable) to both increase the time taken to generate the hash and limiting the speed at which you can brute-force it (bcrypt is much better at this than PBKDF2) and to make it harder to do if you do not know how many iterations to use. Also, because of the rainbow tables we mentioned before, each of these algorithms includes a "salt" value, which is unique to the generated hash. This ensures you can't just build a giant database of hashes.

😈 Your framework should have a way to hash passwords built in. For example, Laravel's is Hash::make. It defaults to bcrypt with Argon2 as an option.

For a brief introduction to password cracking, check out this article by Matt Miller on Beyond Trust. You can also hop on over to TryHackMe's Crack the Hash room to try out password hash cracking for yourself to learn how to defend against it.

Man looking pretty stressed at a laptop
mapo_japan @ Shutterstock #2490729947

This is a thing we've been fighting for a long time. To forget things is human. In an ideal world (where passwords have to exist) everyone would be using a password manager, and they would never ever forget their password. And they'd always have access to that password manager. And they'd always be able to tell their loved ones how to access that password manager in the event of an accident.

😈 You're going to hear a lot about humans forgetting things across this series.

Okay, since none of those things are guaranteed out there in the world, we need to have a way for the user to get in if they don't have access to their password. How do we securely(ish) verify they are who they say they are when they request a password reset?

The most common way to handle this is by sending a one-time link or code to the email address or the phone number they gave us when they signed up that they can then click to prove they have control of that communications channel. We also generally assume that if the user's email address has been compromised, they have way bigger problems than a password reset on our service, so we'll let you in if you possess that one-time code. But! Critically, this requires us to have collected those communication channels up front. Which requires us to store more information about the user. This is why the username also tends to be the email address in services (but this is, IMO, a bad user experience because changing it often becomes problematic).

Related point, we could also send that code to SMS — most folks have a phone after all. But this is a fairly weak form of authentication because of sim swapping attacks. A sim swap, briefly, is where someone calls your phone company, pretends to be you, and then gets the phone company to activate their phone with your number. You lose access to your phone number, they gain it, and can intercept these kinds of codes (this is also true of multifactor). CISA recommends only using SMS or voice as a last-resort option for MFA, and that holds true for forgotten password requests (for the same reasons).

So, assuming we're going to send a reset request to some other communication channel for the user, what do we send them? Well, in order to be secure, it should have the following properties:

  1. Impossible to enumerate in a reasonable amount of time.
    1. This means that you should rate-limit the number of requests to the confirmation page.
    2. It should also be longer than a few characters, if you're using a PIN, at least 6 digits.
  2. Should only be valid for a short period of time.
    1. Generally between 5 and 30 minutes is acceptable. Use the shortest period of time that is reasonable for your user base.
  3. Should only be usable a single time. Once a reset code / pin / URL has been used successfully, it should be invalid for subsequent requests.

For more reading on the topic, I recommend OWASP's Forgot Password Cheat Sheet.

We should talk about multifactor authentication

Permalink to “We should talk about multifactor authentication”

Multifactor Authentication (MFA) is a large topic which I'll cover in more detail in a different post, but for its worth mentioning here: MFA both increases account security and the complexity of ensuring legitimate users don't lose access to their accounts.

So, stay tuned for the deeper dive on how to implement MFA in another post.

One thing that Banks in particular do that drives me to irrational anger is to ask "security" questions, like "What was the make and model of your first car?"

These questions are trying to give you something easy to remember as either a recovery method or a MFA factor, but they invariably ask things that are easy to find out in public or social engineer. Like, anyone can figure out what high school I went to. So instead of answering these in a memorable way, I use a password manager to generate random words to fill out the answers. That way you can't guess them, and I have a shot at remembering them (but otherwise just use my password manager). It's pointless, just give me a proper MFA factor and call it a day.

Stop using these, they're a bad idea.

Just ask NIST:

Verifiers and CSPs SHALL NOT prompt subscribers to use knowledge-based authentication (KBA) (e.g., “What was the name of your first pet?”) or security questions when choosing passwords. - NIST SP-800-63B

There's an old security requirement still floating around (especially in corporate settings) that you must force a password reset every {n} days. This is no longer a good requirement, and you should not force the user to change their password. This has been true for a long time now.

Verifiers SHOULD NOT require memorized secrets to be changed arbitrarily (e.g., periodically). However, verifiers SHALL force a change if there is evidence of compromise of the authenticator. - NIST SP-800-63B

To summarize, your baseline level of secure password storage:

  1. Do not store passwords in plaintext. Use Argon2id, scrypt, bcrypt, or PBKDF2. If you're using a framework it's probably doing this for you, but understand what the framework is doing and which hashing mechanism it has chosen.
    1. If you're going for a NIST certification, that will determine which algorithm you must use.
  2. Consider how you're going to reset the user's passwords when they lose access. This requires collecting an email, phone number, or some other mechanism up front.
  3. Resetting the password should require possession of something else (email, SMS, backup tokens).
  4. MFA adds additional security but also additional account lockout considerations.
  5. Please for the love of the gods do not ask for "security questions".

Alright! So we're getting close to the end of the topic for this post, but what happens if you want to log the user out. You've got a bunch of questions you should answer.

When the user clicks a "Log Out" button, you should do the following things:

  1. Delete their session storage (the cookie and any data you've got on the session server-side)
  2. Redirect them back to a public page.

You do not need to do things like remove their history, because your server should redirect to an authentication page if the session is missing. Most frameworks encapsulate best practices into a logout() method somewhere on your user or session model which will do these things for you.

As a matter of practice, you should offer your users the ability to delete their account. How you implement this depends on local regulations: you may need to anonymize the user's data and remove personal identifiers (but continue to store logs based on retention requirements - this is common in health care settings). However, if you do not have retention requirements, you should remove all related user data from your database that you do not require.

In practice, most applications will need to store a "stub" of the user record. For example, in an eCommerce application you'll want to retain archival records for your orders (which may be synced to other systems), but "what data do we need to retain" is a huge topic and industry-dependent.

Regardless, read up on GDPR or CCPA under the right to be forgotten for further guidelines on how you can delete user data.

Okay, this is getting long, and I'm sure there are some things I missed. If there are bits that you'd like to see me touch on I can update this post later, just leave me a comment below or ping me on Mastodon or Bluesky.

😈 Is this too long? Should it have been a two parter? Let me know.

Part 2 will be on Multifactor Authentication and the various methods for how that works, along with the why, how, and what to do when you drop your phone in a river.

Auth Series 2: Multifactor Authentication

Logos of Authy, Yubico, and a Fingerprint

Welcome to part 2 of the rambling journey through how to do authentication systems yourself (or at least understand how they work). This time it's an extension of part 1, where we learned how to do Username and Password Authentication. It's a topic that's pretty closely attached to that "now that I can log users in, what can I do to make their account resistant to nefarious multidimensional entities who wish to look at their weird game history?"

Okay, maybe you're not asking that exact question, but you do want to make accounts more secure. How do you do that? One way is by implementing a Multifactor authentication system, which is a fairly deep topic. It'll take up the entirety of this post.

I'll be linking off to various resources as we chat so you can find more information, but for this one I'll be quoting NIST and CISA several times because they have solid advice.

There's an adage in authentication circles, the Three Authentication Factors: Something you know, Something you have, and Something you are. (There are "more" sort-of. See this great StackExchange question).

😈 If you're feeling snarky: Something you'll forget, something you'll lose, something that can be impersonated. We'll get into that later in the post.

The first post in this series covered the "something you know" portion of Authentication, the humble password. Today we'll hit both the "something you have" and "something you are" portions. Combining two or more of the three factors is what makes it multifactor authentication.

That's it, that's all there is to it.

Aside: Just how many factors are we going to combine?

Permalink to “Aside: Just how many factors are we going to combine?”

For most commercial applications this is usually "two". One thing you want to do is balance how difficult it is to log in for the user with their willingness to jump through hoops to identify themselves. eCommerce, for example, naturally wants to reduce friction between you clicking on that shiny thing your lizard brain wants and getting through the checkout process to give up your money. I'm sure many of you have had just this conversation with stakeholders. There's not a "naturally correct" answer here. Like many things it really depends on your application and the level of security you want to foster.

A stock image of "authenticators"
SuPatMaN @ Shutterstock #2495179419

We're going to hit on the most common types of MFA for this post, in the order of least-to-most complicated to implement, and I'll talk about the implementation details in each section. Please note, some of the factors I'll mention are discouraged, and I'll put a big ☢️ symbol at the front of their section if you need to be wary of them along with why. I include them because sometimes you just can't get around it due to other factors, but I want to be very clear that there are better options.

Okay, that out of the way, let's go.

We talked about these in the context of a Forgot Password flow in the first post, but One-Time passcodes can also be used in a MFA context when you're sending them to some out-of-bands communication method. The main ways to do this are:

  • Email
  • ☢️ SMS
  • ☢️ Voice Calls

These all demonstrate possession of something, namely your email address or phone number is something you have. This factor assumes that you retain control over the channel in question.

You can do other things, like literally mailing a code to someone via physical mail (which is something the American Social Security office used to do and might still do), but for 99% of applications that's going to be incredibly impractical.

All the rules that apply for forgot password from the first post also apply to generating one-time passcodes, which I'll repeat here:

  1. Impossible to enumerate in a reasonable amount of time.
    1. This means that you should rate-limit the number of requests to the confirmation page.
    2. It should also be longer than a few characters, if you're using a PIN, at least 6 digits.
  2. Should only be valid for a short period of time.
    1. Generally between 5 and 30 minutes is acceptable. Use the shortest period of time that is reasonable for your user base.
  3. Should only be usable a single time. Once a reset code / pin / URL has been used successfully, it should be invalid for subsequent requests.

Then, you send that generated, expiring code to the channel the user has previously provided allowing them to enter the code to finish authenticating their session.

😈 You may also generate a token that you include in a URL parameter so that the user just needs to click to access. This is also how "magic login links" work

Okay, so why did I put the ☢️ symbol on SMS and voice calls? Well, CISA recommends only using SMS or voice as a last-resort option for MFA because of the rise in prevalence of SIM swapping attacks. That, very basically, is where someone calls your phone company, pretends to be you, and gets them to assign your phone number to a phone the attackers control. Here's a Lifehacker Article on the topic.

Anyhow, that leaves "email" as a viable channel, and the general consensus I've seen (this is my professional observation) is "if your email's been compromised you've got bigger problems". Which is both true and disconcerting.

So, if SMS is so bad, why do banks keep using it? I'd love for banks to stop using SMS for their MFA setups by default, but it mostly comes down (I believe) to access vs risk. They assume the majority of their customers will not fall victim to a SIM Swap and the increased complexity of something other than SMS isn't worth the hassle. Banks also have other protections in place to limit the blast radius of a leaked customer account (withdrawal limits, reconciliation checks, a dispute process in meat space, etc).

I don't work in the banking industry though, so that's all just informed speculation on my part. If anyone does and can provide an answer, please leave a comment below.

OTP Codes are able to be phished by an attacker. This is why codes you receive always have messages like no Cthonicbank employee will never ask you for this token, swear in blood that you won't give it out.

This is another "have" factor, in that you possess something that will generate a (usually) 6-digit code which changes every 30 seconds (though this is configurable if both sides know the interval). Most people will know this as the "Google Authenticator" method, but some of us old folks will remember physical RSA tokens (or the one Blizzard did in the earlier days of WoW). Now, there are some major security differences between modern TOTP approaches and the old tokens (namely, those tokens were not based on the modern TOTP algorithm) but the general idea is the same: the server and the client agree on a shared secret key which then feeds a random number generator. The random number generator uses that key (as a seed) and the current 30-second time slice to generate the code. If either side leaks the seed, you'd be able to guess the random numbers.

TOTP is based on an internet standard which you may read all about in RFC 6238, but Wikipedia outlines the basic algorithm.

To gloss over some complexity (if you want to implement this yourself, definitely read the RFC) this algorithm requires the server and the client (Google Authenticator, Authy, a physical token, etc) to establish a "shared secret" during the enrollment process. The key is usually a base32 encoded random string which is then passed into the HMAC-SHA1 (usually, but this is configurable) algorithm as part of the generation process. You also have to agree upon how often the codes are generated and a parameter to tell it how many times it should cycle through the algorithm before landing on a given number. This is all to make it impractical to guess a given number since you'd need to know all of those things (though there are defaults).

This blog post by Øyvind Stegard is a simple implementation of the algorithm in shell scripts, and it'll give you an understanding of the process end-to-end. The algorithm itself is actually quite straightforward (though it builds upon several concepts, like "what the hell is HMAC?").

So! When you enroll an "Authenticator" into a service, most services either provide you a QR Code to scan with all that information encoded, or a string that you should copy and paste into your authenticator. Then, you finish the enrollment by entering the next valid key for the time frame. Checking that token is extremely important because if you do not, you could lock the user out of their account. You do not want that. Always verify that the user can generate a valid code from your TOTP implementation.

😈 This is one of those things where there's very likely a library in your programming language that does this. I don't recommend rolling your own in production, but as far as authentication algorithms go this is one of the easier ones.

To summarize, the process for TOTP Registration is:

  1. Establish a Base32 encoded shared key during registration from the server and share that with the client along with any configurable options (like the token validity interval, number of cycles, hashing algorithm, etc).
    1. The server stores this shared key somewhere alongside the user record.
  2. The client stores that information and starts generating codes.
  3. The user must input a code from the authenticator before you mark the account as having 2FA turned on.
  4. The server validates any generated codes with the shared key from step 1 any time the user needs to input their 2FA token.

Part 2 of the Authentication Examples repo has a page which will let you enroll your TOTP token in an authenticator of your source and generate tokens. It uses a couple of PHP libraries to do this: symfony/lock and spomky-labs/otphp.

Just like the OTP method before, these can be phished, but since their duration is very short this kind of attack often relies on credentials being relayed in real-time through a proxy or some other mechanism for tricking the user to entering the code into a malicious page in real time.

Yubikeys! Or, hardware-based authenticators with WebAuthn

Permalink to “Yubikeys! Or, hardware-based authenticators with WebAuthn”


BestForBest @ Shutterstock #2395187395

Now we're getting into one of the more secure "have" factors. There are several hardware based authenticators on the market which support something called WebAuthn. WebAuthn is the "Web Authentication API" specification which enables using all kinds of hardware devices to authenticate with websites. It's in the same category of thing as a "smart card" that those of you in the government might have used to log into a physical computer. This can also be something like your phone, or a key stored in Bitwarden, but for simplicity’s sake I'm just going to talk about hardware keys in this section.

😈 It's also the basis for using Passkeys, which we'll be talking about at length in a future post.

I'm a big fan of WebAuthn, it's very neat when it's used as a second factor.

For the moment, I'm going to focus on FIDO U2F and FIDO2 (which covers more use cases and is tied up in the explanation of passkeys) but all of these things ultimately run through WebAuthn. Basically, the more modern process for using a hardware device all runs through the same API as passkeys, but for 2FA you simply use this key as a second factor rather than the primary factor.

How you do it winds up getting a little complex (okay maybe a lot complex), but I'll try to summarize the process from the server side of things as succinctly as possible. Let's start with the example from webauthn.guide:

const publicKeyCredentialCreationOptions = {
challenge: Uint8Array.from(
randomStringFromServer, c => c.charCodeAt(0)),
rp: {
name: "Duo Security",
id: "duosecurity.com",
},
user: {
id: Uint8Array.from(
"UZSL85T9AFC", c => c.charCodeAt(0)),
name: "lee@webauthn.guide",
displayName: "Lee",
},
pubKeyCredParams: [{alg: -7, type: "public-key"}],
authenticatorSelection: {
authenticatorAttachment: "cross-platform",
},
timeout: 60000,
attestation: "direct"
};

const credential = await navigator.credentials.create({
publicKey: publicKeyCredentialCreationOptions
});

First, you need to establish the website as a Relying Party (RP - this term will come up again in the post on OIDC / SAML). This kind of second factor is tied explicitly to the website in question so that you can't use common browser hijacking techniques to trick the user into giving you a valid key (unlike TOTP codes). In this example, the website is duosecurity.com and the name is Duo Security. The name can be whatever you want, but the id must match the site that you're presently authenticating to.

The second bit is establishing the user stanza. This is where things get a little weird. Let's look at it again:

user: {
id: Uint8Array.from(
"UZSL85T9AFC", c => c.charCodeAt(0)),
name: "lee@webauthn.guide",
displayName: "Lee",
},

The thing that probably sticks out to you is that id stanza. Where did they get UZSL85T9AFC? Why is it a Uint8Array? Yeah, you'll have to go to the spec for that:

The user handle of the user account. A user handle is an opaque byte sequence with a maximum size of 64 bytes, and is not meant to be displayed to the user.
To ensure secure operation, authentication and authorization decisions MUST be made on the basis of this id member, not the displayName nor name members. See Section 6.1 of [RFC8266].
The user handle MUST NOT contain personally identifying information about the user, such as a username or e-mail address; see § 14.6.1 User Handle Contents for details. The user handle MUST NOT be empty.

Okay, that's a mouth(eye?)full, but it's relatively simple. The ID must be a byte sequence (which is why it's a Uint8Array) and it must be opaque, not displayed to the user, and used for authorization decisions. Cool. Why's it seemingly random?

Couple of reasons, but the primary one being some token authenticators will make a discoverable (resident) key, which'll be visible on the device. As the Relying Party you have to both provide this id and store it with the user record. Frequently, I've seen this be the primary key for the User, either the auto_incrementing ID or a UUID, or a generated uniqid(). This is easy to do when you've already got a user record in a MFA context, but we'll talk a lot more about this in the Passkey segment.

Okay, cool. The only other stanza I think worth explaining here is the authenticatorSelection. There are a couple of options you can pass here. I'll cover them at length in the Passkeys post (I'm saying that a lot, aren't I?), but the one I want to call out right now is credentialProtectionPolicy.

credentialProtectionPolicy can be set to userVerificationOptional, userVerificationRequired, or userVerificationOptionalWithCredentialIDList. That last one requires some more explanation, so I'll leave it alone for the moment, but the first two control whether the authenticator is encouraged to ask you for a pin code or a biometric. In the case of something like the Yubikey 5c which has no biometrics, this will be a prompt to enter a pin into a browser-native popup. For things like an iPhone, it'll likely be FaceID. Note that userVerificationOptional doesn't mean that the device won't prompt you for a verification, some devices may always prompt for verification if they so choose. The credProtect extension documentation covers the various scenarios.

For MFA scenarios, this is usually set to "optional" because FIDO 2.0 did it that way, and it's less friction for the user who has already provided a password for the first factor.

😈 Some sites will still prompt for verification in a MFA scenario. I know Cloudflare and Zoho do this.

So, that was a lot of explanation. What happens after you initiate this request? Well, the user will authenticate with their WebAuthn token, and you'll get back a PublicKeyCredential. I want to quote webauthn.guide once again because I find this phrasing very amusing:

After the PublicKeyCredential has been obtained, it is sent to the server for validation. The WebAuthn specification describes a 19-point procedure to validate the registration data; what this looks like will vary depending on the language your server software is written in.

It's a lot. There are libraries that do this for you. Please use the libraries.

Anyhow, once you've validated this on the server side you store the public key you received attached to that userID (and the user record) you got before. The key will also generate an ID for itself, and you'll want to store this alongside your user for non-resident keys. You will want this to be a one-to-many relationship between the User and the Keys, because for physical tokens, this public key isn't portable (Apple and Google store keys in your keychain, as do some password managers, but you cannot assume that the key will be portable).

Okay. We've registered the token, how do we validate it when the user's logging in? From webauthn.guide:

const publicKeyCredentialRequestOptions = {
challenge: Uint8Array.from(
randomStringFromServer, c => c.charCodeAt(0)),
allowCredentials: [{
id: Uint8Array.from(
credentialId, c => c.charCodeAt(0)),
type: 'public-key',
transports: ['usb', 'ble', 'nfc'],
}],
timeout: 60000,
}

const assertion = await navigator.credentials.get({
publicKey: publicKeyCredentialRequestOptions
});

The main bits you'll want to focus on in that example is challenge and allowCredentials.id. The challenge is a Uint8Array that you generate on the server and provide to the credential to sign with its private key. The allowCredentials.id should match the we got back from the key in the registration step. You'll receive a PublicKeyCredential object back from this request, assuming the token is present. If you get an error, you'll want to parse the error and provide a button for the user to try again with the token (or to use an alternative method).

Here's how PublicKeyCredential looks:

PublicKeyCredential {
id: 'ADSUllKQmbqdGtpu4sjseh4cg2TxSvrbcHDTBsv4NSSX9...',
rawId: ArrayBuffer(59),
response: AuthenticatorAssertionResponse {
authenticatorData: ArrayBuffer(191),
clientDataJSON: ArrayBuffer(118),
signature: ArrayBuffer(70),
userHandle: ArrayBuffer(10),
},
type: 'public-key'
}

There's yet more validation you need to do on this response, but ultimately you'll be checking to ensure that signature matches what you get if you use the public key you've stored for the user to sign that same challenge you produced before. If it does, you're good! If it doesn't.... well the key is bad or has been tampered with (if, for example, someone's made a key that responds to all credential requests or something).

Okay. That's it! We did it! WebAuthn 2FA!

😈 That is far from it, dear reader. There will be an entire post about WebAuthn for passkeys (which are basically the same procedure but there's a lot of detail to talk about).

It's worth doing this hands-on if you really want to get the concept, so Google has a great tutorial for doing this end-to-end along with a Glitch site with the code to it located here.

Right, this section is going to be relatively short because it's got a ton of overlap with the WebAuthn section because practical biometric auth for the web follows the same API patterns. Biometrics are the "something you are" factor, and are most commonly a fingerprint or a face identification. It can also encompass retinal scanners and other more... esoteric ideas, but you're likely to only ever interact with Fingerprint or Face.

😈 Genetic Marker testing for web here we come?!? (Gods I shouldn't give terrible people ideas)

Relatedly, on the web you'll likely not know which biometric authentication factor you've encountered or even if it's biometric at all. That's because (as far as I'm aware) the WebAuthn specification has no mechanism for saying "only allow devices with biometric attestation". The best you can do is set credentialProtectionPolicy: "userVerificationRequired" and hope the device in question is biometric. Yay!

Like I mentioned above, Apple and Google are going to use biometrics to authenticate, and some Yubikeys do biometric auth (with fingerprints) like this one.

So there are some other things worth mentioning (because we need to talk about recovery):

  • Pre-Shared Secrets
  • Lookup Secrets
  • Multiple Factor OTP Devices

Let's look at each of these just a little bit.

These are things that are shared between the server and the user. Passwords are a pre-shared secret. However, backup codes are also pre-shared secrets. We'll be talking about those in a minute, for recovery.

These are cool, and totally impractical. This is essentially a generated table of random whatevers which is arranged in a grid, so that you can prompt a user to lookup (for example) the word in column 2 row 24. This table should be random and unique per user.

The only thing I can think of where I've seen anything close to this is 2048 BIP-39 seed phrases which is commonly used for crypto wallets. That spec is just a giant list of possible words, arranged in order for a given cryptographic seed.

So, there's another cool thing you can do with Yubikeys. Yubikey makes a Yubico Authenticator app that lets you use your yubikey as an OTP device. So you have to have the Yubikey in order to generate the OTP. It's neat. I use this for sensitive things that don't support FIDO/WebAuthn.

Phone sitting precariously next to a gutter
monte_a @ Shutterstock #1784028497

Remember how I said that humans tend to forget things? Well, adding "something you have" factors to your authentication flow adds the possibility that they'll lose something in addition to forgetting it. This could be as simple as "I left my yubikey at home" or "I dropped my phone in the river". Yeah, troublesome.

This means you need to provide the user ways to break the glass to get back into their account. There are several ways you can handle this, but let's talk about the most common first. Backup codes!

If you've ever enrolled in a 2FA system yourself, you've likely encountered this. After you enroll the device you'll be presented with some number of "backup codes" in a window that warns you that if you lose your device you'll need those codes to get in. If you don't save them, you'll be locked out forever if you don't have it.

As the server operator, this is as simple as generating random impractical-to-guess codes and storing them attached to the user record. You'd then allow these codes to be used one time each as a bypass for the MFA device. Each time a code is used, remove it from the pool of valid codes you've made for the user.

The user, for their part, needs to store those codes somewhere safe, often printed or written down on a piece of paper. This is why recovery codes are often simply alphanumeric and relatively short (to allow for ease of entry).

You should also periodically remind users about their backup codes, and offer to let them regenerate them if it's been a long time since they were generated to reduce the chance that they lose those too.

Now, many services stop here and will not offer further recovery if the 2FA factors are all lost. Too bad, user, should have kept those codes. This is sensible for many applications, but not great for something like a bank.

Another option you have is a customer service based approach where you have the user use a customer service line where they have to prove their identity to a human in order to have MFA removed from their account.

Designing a customer service protocol around this is beyond the scope of this document (but if you want to talk about it, hop on over to the contact form, and we can set up some consulting time), but you want it to be robust and resistant to impersonation.

😈 Remember kids, customer service attacks are a big reason SMS hijacking is a thing.

Like forgot password considerations, you could allow an SMS or Email code be sent to the user to let them bypass 2FA, but I do not recommend this as it reduces the account to the weakest form of MFA and makes offering OTP / WebAuthn factors.

Okay, that was... a lot? Less words than on usernames and passwords, but still a fair amount of words to cover 2FA.

Stay tuned for part 3, where we'll be talking about single sign on (SSO) and associated protocols (SAML, OAuth/OIDC) along with the how / why you might want to do that.

Lemme know in the comments if there's something I missed or anything else you'd like to see.

Auth Series 3: SSO

Intro image showing the logos of SAML, OAuth, and OIDC

Right! Welcome back to the third post in the Authentication Post series and this time we're going to talk about single sign on (SSO) and other federated identity protocols.

I'll try to keep this post under the word count of the previous two, but like many auth things… this could be a lengthy topic.

In the previous posts in the series we covered how to authenticate users with a username and password on a single service, and how to add multifactor authentication to that account. What happens if you want to create another application? One option you can choose is to simply make a second user base for that application. That might be the right choice if you want those user bases to be completely separated from one another. But what if the two applications you've created are closely linked together? You could copy your existing user database to the new service, or link both services to the same database, but now you're having to manage the complexity of the user experience between two different sites. FIDO/U2F/Passkeys don't work cross-domain either, so you might be dealing with the complexities around that as well.

Plus, the users will have to enter their credentials on both sites any time they want to use them at the same time.

SSO solutions were built to handle these kinds of use-cases, allowing a user to log into multiple sites with a single set of credentials while minimizing the number of interactions the end user needs to perform. There have been many iterations of this over the years, so I'll give you a brief (non-exhaustive) history of some options and where they've gone.

One common option in the "olden days" was to have a single parent domain where you authenticated the user, and then used 3rd party cookies in order to trigger the authentication on multiple target sites by granting those sites access to that cookie. This was really common when you had a root parent domain (example.com) and your additional sites were all subdomains of that first site (service1.example.com, service2.example.com) because it was relatively easy to set a cookie that's readable on all those domains. You'd then wire the backend services to the parent domain for handling account actions (meaning you'd land on example.com to change your password, for example). The good news is you can still do that, so long as you set the domain property of the cookie to that parent domain, subdomains can still read that cookie.

But what if your authentication system is on a different domain as the services (one real life example of this is google.com vs youtube.com and how they do SSO on that is actually pretty fun)? Well, for a long while you could do something like an AJAX request to the backend service and have that service set a cookie that could be read on any domain. That's known as a 3rd party cookie. You might have heard a lot about those, because they're also used for surveillance capitalism and serving you ads. The browser vendors have been slowly blocking them by default for years (though Chrome apparently is walking back their plans because it'd impact their ad revenue). So that technique no longer works for auth because ad vendors are... well... I'll stop there.

😈 Fun fact, when you sign into Google it also issues a 302 redirect to Youtube to sign you in on Youtube to get around the cross-domain problem. (source)

Another fun element of the rise of social media platforms was the rise of using a social account to log into another service. Many services today offer the ability to log into their services using your Google, Facebook, (rip) Twitter, and etc. accounts in lieu of signing up for yet another website. This has some benefits for the end user, you don't need to remember yet another password for the service, you can just click on the "sign in with Google" button, asked to share some information, and then you're logged in! Easy.

The downside, of course, is you're locked into the social provider's whims, and service providers generally only bothered to support the biggest players (this was especially true before OAuth 2 / OIDC).

Anyhow, in this model you can't just rely on setting cookies willy nilly, that'd be insecure as all get out (you don't want every website on the internet having unrestricted access to your Google information — only Google Ads get to do that), so you need some sort of protocol to handle the exchange of information, first to ensure that the site in question is allowed to even ask you for your account information and a mechanism for securely sharing the bits you consent to sending back to the service.

That's OAuth (and now OIDC... I'll explain shortly)!

There are a bunch of other ways you can do authentication against parties you don't control, including:

  • SAML (We'll talk about this)
  • Kerberos
  • LDAP
  • RADIUS

A lot of these are context-dependent. For example, I've personally seen more SAML integrations with Universities than I've seen of the other protocols. Kerberos is pretty common in Enterprises (combined with Active Directory often alongside an LDAP system).

If you own the system, why use a protocol?

Permalink to “If you own the system, why use a protocol?”

So, like I mentioned above, using a combination of cookies and sorcery, you could build your own authentication systems with username and password authentication that works across a number of different properties you control. I'm going to recommend you do not do that. I strongly recommend that if you're going to build a centralized authentication system, you use a protocol (and I'm going to go further and say you use OAuth2 / OIDC). This is because it makes it a lot easier to implement across those sites. There will be libraries that you don't have to write yourself. The pathways are well-known. The cognitive overhead is smaller. If you want to let a third-party into your system you don't have to do something bespoke.

It'll save you a lot of headache after the initial headache of understanding how the heck these protocols work. Walk with me and I'll help you out.

SAML 2.0 - a versatile and complicated protocol

Permalink to “SAML 2.0 - a versatile and complicated protocol”

Carlo Toffolo @ Shutterstock #795590020
Carlo Toffolo @ Shutterstock #795590020

I don't want to spend a lot of time on SAML, but it's near and dear to my heart ever since I spent a full week reverse-engineering SimpleSAMLphp and reading the spec until my mind exploded. There are some similarities to how OAuth and OIDC handle the cross-site communications though, so it's worth discussing.

SAML is the "Security Assertion Markup Language", and by that it means it's XML. It's a lot of XML. Every bit of the communication protocol is sending large blobs of XML back and forth. The way this works is essentially this process for Service Provider (SP) initiated login:

😈 Before any of this can work the service provider and the identity provider must be configured to allow this. Namely, each side of the exchange must set up some configuration options to identify themselves to each other (otherwise any random site on the internet could try to trick you into signing into the identity provider).

Also, I'm going to use "service provider" through the rest of this post, but OAuth tends to call them "Relying Parties" or simply "the client".
  1. The SP, example.com, issues an <samlp:AuthnRequest> to the Identity Provider (IdP) greatlogins.test by a HTTP GET or HTTP POST request in the SAMLRequest param (this XML is deflated and base64 encoded to fit in that param).
    1. There's also an Artifact Binding which uses SOAP and a reference ID to allow a lookup rather than send the whole response but I've literally never seen this implemented in the wild. Not a single time. Have you?
  2. The IdP base64 decodes and inflates the AuthnRequest and validates it contains what it expects. Namely, that the SP is on its Allow List, and that it has been signed (and optionally encrypted) correctly by that SP and the request has not been tampered with.
    1. Yeah, setting up a SAML connection requires sharing certificates ahead of time. SAML Requests are signed, and optionally encrypted.
  3. The user is shown a login screen where they enter whatever credentials they need to (the SP doesn't need to care how this happens).
  4. The user is redirected back to the SP (again via a GET or a POST) which contains a <samlp:AuthnResponse> in the (you guessed it) SAMLResponse param. This too is deflated and base64 encoded.
  5. The SP validates the response in the same way the IdP did (ensuring the response has not been tampered with) and then pulls attributes out of the response to make authentication / authorization decisions based on that. These attributes are configurable, and are usually a thing you want to negotiate when doing the setup process.
😈 Why did I use a .test there? Because .test is a reserved TLD and there's no real website that could resolve to! (Like example.com)

Did I make that sound easy? Well there are a lot of things that can go wrong, but the good news is those are relatively predictable if you know how XMLSec works... Right. Yeah. Okay, so tl;dr it's usually a problem with configuring the shared secrets or missing response attributes. It could be other things, but it's there.

One major thing you're going to run into is there still aren't a lot of deep tutorials (that I'm aware of) about SAML. Here's a decent one.

That's all I want to say about SAML right now, but I did want to include it so you can see the similarities in...

Also known as "the precious". This is my most preferred SSO protocol these days, not least of which is because of breadth and depth. There are a number of software packages that support these two protocols. Doing simple implementations is very straightforward, but there's a lot of configuration and security features. There's even considerations for signing in to smart devices that don't have a web browser. There are definitely rough edges, but, it's the one that comes close to being the best we've got for the most people.

Like I mentioned in the history lesson, OAuth was originally created to allow users to grant 3rd parties access to a social account in order to do things on their behalf. This varied, but it was usually in the context of "allow this site to post as you on social media" to enable various kinds of experiences. While OAuth by itself grants authorization rather than authentication, a lot of folks were using it for both. Namely, if you're issued a token that can act on behalf of a social account, isn't that proof enough of authentication?

😈 No! But people were using it that way anyhow.

That's where OIDC comes into the picture. OIDC (or Open ID Connect) provides a series of protocols for authentication and it's often paired with OAuth 2.0 for authorization. This, I think, is why you hear folks use the terms basically interchangeably these days. I've been caught doing that very thing because for most audiences the distinction doesn't matter.

That out of the way, here's how a typical flow with OIDC works (try to spot the similarities to SAML). Now, there are several OAuth grant types and flows that you can and should be using which I will get into, but we'll start with the simplest one that's intended for end users: response_type=token (which is the deprecated Implicit grant). You should not use this grant type any longer, but it's the simplest version, and we'll use it to build your knowledge.

😈 The actual simplest one is client_credentials which involves sending a client_id and client_secret in exchange for a token. It's intended for machine-to-machine use cases, not for end users.

Like SAML, the IdP must be configured with knowledge of the SP (in the form of an "Application") which grants the SP a client_id and the SP tells the IdP what URLs it's allowed / expected to redirect you to.

😈 There's such a thing as client autoregistration, which is a thing that Mastodon uses, which allows you to just create a client on the fly. I... don't really like it but in some cases it's necessary. Like Mastodon.

The Implicit flow works like this:

  1. The SP redirects the user to the IdP's login endpoint (optionally autodiscovering this URL by querying /.well-known/openid-configuration on the IdP domain), including the client_id, response_type (token for this example), scopes (what information you want to request), state which is an identifier for your application to manage its "state", and the redirect_uri.
  2. The user is prompted to log into the IdP. Assume they do so.
  3. The user is shown a screen that shows the application that is requesting the user's information, along with an enumerated list of what information the application is requesting.
  4. Assuming the user says "yes go go go", the IdP redirects to the redirect_uri (assuming it's on the allowed list) along with a token parameter (and a refresh token to get a new one when that token expires).
  5. The Application can just use the token to then get information about the user from the IdP. Simple!

Now, the implicit flow is fairly vulnerable to access token leakage and token replay attacks, and they can't be bound to a given client. Instead, what you should be using is the Authorization Code flow, with PKCE. Let's build on the Implicit flow and add Authorization Code.

Here's how that looks:

  1. The SP redirects to the IdP in the same way as in the implicit flow, but instead response_type=authorization_code.
  2. Same
  3. Same
  4. The IdP redirects back with an authorization_code which is short-lived (usually 30 seconds).
  5. The SP makes a call to the token endpoint with that authorization_code from the backend along with its client_secret. Since an attacker won't have that secret, they cannot exchange the code for a token, preventing some leakage issues. The timeout prevents replay attacks.
    1. You can configure a public client that doesn't need the secret to do the exchange.
  6. The IdP returns the token to the SP, and marks the authorization code as used (though some providers allow you to configure how many times the code's valid to prevent weird race condition problems).

You might be noticing an issue here, if you're familiar with mobile apps. If you're authenticating directly from a device, the client must be public. There's no way to keep a secret a secret in a mobile app. Someone could decompile it, or otherwise coerce the information from the device (This is also true of Single Page Apps). So, how do we make those calls a bit more secure? PKCE!

😈 Sidebar: PKCE is also useful for apps with client secrets to prevent CSRF attacks.

Let's build upon our previous flow.

  1. The response type remains the same, but in this case the SP creates a code_verifier which to quote oauth.com: "This is a cryptographically random string using the characters A-Z, a-z, 0-9, and the punctuation characters -._~ (hyphen, period, underscore, and tilde), between 43 and 128 characters long."
  2. The SP makes a SHA-256 hash of the code_verifier which is then base64 encoded and sent along with the request. This is the code_challenge
  3. Proceed along until step 5 of the previous step - this time, instead of sending just the code, you also send the code_verifier along with the request.
  4. The IdP uses the code_verifier to generate a code_challenge and checks to see if that matches the code_challenge it received in step #2. If it does, it can be assured that the request hasn't been intercepted partway through since only the client had that information.
  5. The IdP returns tokens as established.

And that's PKCE in a nutshell.

vector zefirka @ Shutterstock #1936423573
vector zefirka @ Shutterstock #1936423573

Right, so I've said "the IdP returns tokens" a number of times now, and while that's accurate there's some more detail that we need to cover. What is a token?

It could be several things, turns out, but at it's most basic in OAuth parlance it's a "Bearer" token which allows access to resources on a 3rd party when sent along with the request. A lot of services will return an opaque token which is just an internal identifier to the IdP and means nothing outside of that. It could also be a JSON Web Token (JWT) that is an encoded and (usually) signed token that contains information about said token.

Here's an example of a (unsigned) JWT Access token from Auth0:

{
"iss": "https://my-domain.auth0.com/",
"sub": "auth0|123456",
"aud": [
"https://example.com/health-api",
"https://my-domain.auth0.com/userinfo"
],
"azp": "my_client_id",
"exp": 1311281970,
"iat": 1311280970,
"scope": "openid profile read:patients read:admin"
}
😈 For Access tokens Auth0 recommends treating them as opaque always, regardless of format. That is to say they expect you to call endpoints with the token, not introspect it.

There are also refresh tokens, which are how you keep a user logged in over an extended period of time by exchanging them for new access tokens (and ideally rotating them) and ID Tokens.

Let's look at the latter token.

It's worth calling out at this point that OIDC endpoints (remember how I mentioned that OIDC is concerned with authentication while OAuth is concerned with authorization?) can also return an id_token which will contain information about the user along with a series of claims (depending on which scopes you send along with the initial request).

In order to get this token, you have to pass the openid scope in the first request, and it'll return the id_token along with the access and refresh tokens. The id_token is always a JWT, and it should be cryptographically signed by the JSON Web Key (JWK) that's present in the /.well-known/openid-configuration endpoint I mentioned above.

Here's an example Auth0 ID Token (without the signature bits):

{
"iss": "http://my-domain.auth0.com",
"sub": "auth0|123456",
"aud": "my_client_id",
"exp": 1311281970,
"iat": 1311280970,
"name": "Jane Doe",
"given_name": "Jane",
"family_name": "Doe",
"gender": "female",
"birthdate": "0000-10-31",
"email": "janedoe@example.com",
"picture": "http://example.com/janedoe/me.jpg"
}

You'll notice that there's a fair amount of information about the user in encoded in that token, but there are a couple of things I want to call out. First, everything from "name" onwards in that example isn't covered by the registered claim names from the spec. They're custom values Auth0 has added based on the scopes you're asking for. Second, the registered claims have special meanings which you'll want to take account of.

  • iss is the issuer field, this should be the IdP.
  • sub is the user's unique identifier, this will usually be an auto-incrementing id, or a UUID, or another unique id.
  • aud is the intended "audience" of the token, this is usually the client id of the app that requested the token.
  • exp and iat are the expiration time and issued at time, respectively in unixtimestamp format.

You can use any of the information in the token inside your application for displaying to the user or making local assertions but it is critically important that you validate the claims in the token have not been tampered with. Especially ensure that the iss and the aud claims are what you expect them to be and that the exp claim is in the future. It's not a huge deal if you're using ID tokens just for user data, but if you use them in lieu of access tokens (you shouldn't, more in a second), token forgery becomes a big problem.

For more information on validating ID tokens vs Access Tokens, please have a read of this excellent Auth0 Article

But is it really bad to use an ID token as an access token?

Permalink to “But is it really bad to use an ID token as an access token?”

Right. Okay. So there are, let's say, different opinions on this subject. Some folks, like Google think it's alright to use the claims in an ID token to establish the user's identity and use that as a proof of Authorization. If you have a highly coupled frontend and backend, this can be "okay" in that your aud claim should be something you generally expect from the backend.

Auth0 agrees with this statement in the article I linked, stating:

As said above, an ID token proves that a user has been authenticated. In a first-party scenario, i.e., in a scenario where the client and the API are both controlled by you, you may decide that your ID token is good to make authorization decisions: maybe all you need to know is the user identity.

If that's enough for you, so be it! But please for the love of the gods validate that the token isn't forged or expired. You still want to have a list of allowed aud claims as well, don't just allow any aud to pass.

😈 So there's also a cool thing called Demonstration of Proof-of-Possession (DPoP) which uses light cryptography to bind a JWT access token to a particular device. It exists so that if an access token leaks it's useless to an attacker since they don't have a private key needed to sign requests.

Relatedly, ID tokens lack scopes, so you won't be able to rely on them for granular authorization from the IdP. Meaning your backend will need to make authorization decisions some other way. You could configure the IdP to pass those as custom claims in the ID token, but at that point just use the access token, that's what it's for.

You're probably getting sick of hearing about tokens, but there are a few more things I should point out if you're digging into the world of OAuth/OIDC and want to avoid some footguns.

  1. Access tokens should be short-lived. This limits the damage should one be stolen. Auth0 defaults the lifetime to 24 hours. I'd recommend far less than that unless your application needs longer access token lifetimes.
    1. JWTs Cannot be revoked, by the way, so if a long-lived JWT access token leaks... you've got a problem on your hands.
  2. If you issue refresh tokens (via the offline_access scope), enable refresh token rotation. This way you can't reuse old refresh tokens.
  3. Limit the lifetime of the refresh token. This can be much longer than the access token, but it should be directly related to how long you want a user to be logged into your application. 7 days, 30 days, or the like are good ideas.
  4. Always validate your JWTs fully. There are libraries for this.
  5. Don't store sensitive data in ID Tokens. They're trivially decodable.
  6. Do not attempt to store client_secrets in a mobile app or a SPA. They will leak.

Right, so I've explained at length the how and the why of OAuth/OIDC, but none of that information actually solves for the "I want to silently log you in to multiple applications". You still need some sort of initiator for this to occur.

This is where OIDC silent login comes into play. You see, most IdPs support this protocol which allows you to skip the login prompt on your IdP if the user has logged in recently. How long between credential prompts is generally configurable, but usually I've seen it set to "the browser session". The way you do this is passing prompt=none when doing the initial redirect to the IdP. If the user's logged in already, the application just redirects them straight back to your site and you proceed with the token exchange as usual.

If they're not, then they're shown the authentication prompt as normal.

SSO Achieved! ... well, not quite. The user, in this scenario, still has to click "login". This is usually okay, since the process winds up being "click login, suddenly the app knows who I am", but sometimes that's unacceptable.

From that point, you have a couple of options:

  1. Combine this technique with the shared cookie thing from above where the cookie stores a trigger to auto-activate the login flow
  2. Redirect around your sites in a loop, hitting the authorization endpoint on each of them
    1. That can get unwieldy when you've got a lot of sites.
  3. Do what StackExchange did with LocalStorage on a shared domain.
  4. Cry, just a little. It's cathartic.

Relatedly, this is also a problem with logout.

Everything that applies to SSO also applies to SLO and the techniques also apply there. OIDC provides a mechanism for a redirect after hitting the logout endpoint, and you can use that to redirect them around and immediately end the session on all sites.

One other thing you can do is have a short-lived access token and revoke the refresh tokens on user logout. That'll ensure that when the access token expires they'll be signed out. So...yeah!

Likewise, if you're using that centralized cookie or local storage approach, you can just destroy those storage bits and force a logout on that end.

Well, I failed at making this post shorter than part 2. We'll try again in Part 4: Passkeys! Lots and lots of information about Passkeys.

So. Many. Passkeys.

Auth Series 4: Passkeys

Auth Series 4: Passkeys
I'm apparently obliged to tell you: The passkey icon is a trademark of FIDO Alliance, Inc

Here we are again! This time we're going to talk about one of my favorite subsets of this topic: Passkeys. You should mentally prepare yourself for this, it's going to be a lot of information and a fair amount of it is dense. I'll be using information from my other post on Passkeys so you don't need to read it first, but you can if you want.

😈 Guess what! "Passkey" doesn't actually have a real definition, it's a marketing term that a bunch of folks vied for control of. I think we've settled on a common definition, but keep that in mind.

That's as good a place to start as any, so let's gooooooo.

Okay, so, "passkeys" are rooted in the idea that passwords are bad, and we should securely eliminate them. As a reminder, passwords are "bad" because:

  • You have to remember them
  • People tend to reuse them across multiple websites so that they can remember them
  • Humans also tend to create insecure passwords (hunter2 comes to mind) that are easy to guess
  • Which means if they leak, the blast radius can be real big for a given individual
  • They can be phished relatively easily

The best advice that I can give to an individual in the absence of passkey technology is to get a password manager and have that autogenerate and fill secure passwords for you. Apple and Google both have password managers built in, but my preference would be Bitwarden or 1Password. This is going to be good advice until passkeys become reliably ubiquitous, if they ever do.

😈 I think they will, but if you want a more pessimistic view read this article.

Viewed with that lens, you want something that can be used in place of passwords that doesn't have the same security issues as passwords, but can accurately attest the identity of the user signing in to the service. That's what passkeys are. Technically, passkeys are a cryptographic private/public key pair tied to a particular service (they cannot be reused across services) which are stored somewhere that hopefully only the unique user in question has access to. This can be a bunch of things, which we'll talk about in a minute, but they're often also behind a biometric authentication mechanism as well. If you remember the second post in this series, this will be a combination of "something you have" and "something you are".

😈 BTW I'm not actually using the prevailing definition of passkeys here, I'll correct that later in the post. Stick with it.

Passkeys use asymmetric (public-key) cryptography to do their magic, so if you're unfamiliar, the Wikipedia page is a great resource to learn more. But, basically, when enrolling a passkey your device will generate a private key and a public key. The public key is given to the service and stored there. Because of how public key cryptography works, the public key and private key can each generate messages (signatures, or encryption) that you need the other side to verify. The private key stays with the user, the public key goes to the service.

So, when you as an end user start using passkeys to log into things, you'll begin to collect a pile of key pairs stored ... somewhere. Right, okay. So now we need to talk about the various things that can generate/store passkeys. I'll start from the end-user perspective and then move over to the developer perspective after that.

There's going to be a lot of "it depends" here. Also, a lot of "it's being worked on". I'll go ahead and say it now, the UX around passkeys still needs a lot of work (and folks are working on it).

This is probably the most confusing thing to the average user, I think, because the answer is "wide" but also the major platforms are trying to make it seamless to people who don't think about authentication all day every day (... I don't have a problem). So, the options you have right now are:

  • Apple Password Manager with Apple biometrics (Face ID or Fingerprints)
  • Google Password Manager with whatever biometrics your device supports
  • Windows Hello on certain devices (you need a secure element on the device, most modern machines will have this) with a pin or a biometric
  • Physical security keys, YubiKey being the most prominent, with a pin or a biometric (YubiKey sells a fingerprint reader fob)
  • Certain Password Managers like Bitwarden, KeePassX, 1Password
    • This requires an app or a browser plugin to be installed

Note that I listed the platform-specific ones up front, because that's where the "seamless migration" stuff is happening. Apple, Google, and Microsoft are all positioned to make the downsides of passkeys a bit less stabby to the end user. The other options require more effort and understanding to function...and there are some other sharp edges that I'll talk about shortly.

Anyhow, regardless of which option you use, there's a key pair associated with each service being generated on that device. What the prompt looks like when you generate a passkey depends a lot on the platform. Here's what my iPhone looks like when trying to register a new passkey on passkeys.io:

iPhone showing both Bitwarden and Passwords App
😈 This is actually a lot better than it used to be, because it now tells you where the passkeys get synced.

Because I have a saved passkey in my iCloud passwords, I can also use that same passkey on the mac. Here's how that UX looks:
Showing the same passkey on the Mac

So how did that work? Well, for Apple Password Manager with iCloud (and Google Passwords, and the password managers), when you generate a passkey it's synced to every device that's signed in to your iCloud account via the intertubes along with the rest of your password vault. All of that is encrypted with your device code so if someone manages to steal the backups they can't decrypt the keys (hopefully). If they didn't sync the passkeys across devices you wouldn't be able to access your account unless you had the specific phone or computer you used to create the passkey on. It'd also be a major problem if you lose your phone, because that passkey would be unrecoverable.

This is especially true of physical security keys. The private key, by design, cannot leave that device. If you lose the device, you lose the passkey. You might remember that constraint from the Multifactor post.

But wait, you might be wondering: What happens if I generate a bunch of passkeys in Apple's password manager, and then I decide to leave the Apple ecosystem? "We're working on it" is the answer to that question. The FIDO Alliance has released draft specifications for Secure Credential Exchange so you can migrate from one provider to another. At the time of this writing, though, the answer is "you don't". You'd have to go to each service and change out your passkey.

Passkey support is still a bit spotty in some browsers. Librewolf, for example, only reports partial passkey support so when browsers try to detect what UX it can show you, it fails and falls back to just showing a "please touch your device" popup with no further explanation. You could conceivably plug in a USB key after the prompt happens, but the messaging you get around that is very bad since you're now in the "edge case".

Right, fine, but how are Passkeys more secure?

Permalink to “Right, fine, but how are Passkeys more secure?”

Well, the first thing is you don't need to remember anything about them other than you used one to sign in to a service (and they can be autodiscovered, sometimes, we'll talk about that in the dev section). The second thing is they can't be phished. How?

Remember earlier I mentioned that passkeys are bound to a particular website / service? There are a couple of ways this is accomplished (which has a lot of overlap with the multifactor post on WebAuthn) is with the declaration of the Relying Party, the way browsers have implemented the communication channel, and the fact that those attestations are part of the authentication ceremony. While it's theoretically possible the FIDO alliance missed something critical, there are strong assurances built into the spec that the service prompting you for the passkey is the same service you registered with.

The final bit is stealing private keys is also difficult, depending on the storage mechanism. For password managers, it's as secure as your password service, typically encrypted with a thing you, the human, know. For Physical keys, the key cannot be extracted barring exotic attacks and physical access to the key.

😈 There was a fun Yubikey side channel discovered earlier this year, but it also requires knowing your pin or a valid service login.

So, in summary, passkeys are better than passwords because:

  • You can't tell someone a passkey, or have it stolen from a service's database.
  • It can't be phished, since it doesn't ever leave the device in question.
  • If you're using a syncing service, you also can't really lose them since they're backed up to an account.
    • But you could drop your phone and break it, or lose a YubiKey or the like.
  • Assuming the place storing the private key portion of the passkey is secure, it's quite difficult to steal them.

The Technical bits, how do you implement Passkey Authentication?

Permalink to “The Technical bits, how do you implement Passkey Authentication?”

Right. So here's where that "definition" of passkeys gets in the way of explaining this. There are two things that could conceivably be called passkeys: Resident and non-resident keys (also known as discoverable and non-discoverable).

Resident keys are generated when you register with a service and are stored on the device in question. Your browser (or other integrated software) can query the device for these keys to figure out if you've registered with a service. This allows you to use only the passkey to authenticate, no need for a service identifier like a username. In short, Resident keys can replace both usernames and passwords.

Non-resident keys use an identifier held by the service to deterministically generate the public / private key pair repeatably. This means the key in question does not need to store any information about the service its registered with, because it can just regenerate the key from that service identifier each time. This means that you still need a username or other account identifier to figure out what you need to send the key for it to generate the proper attestations. You cannot query the device for what services this has been done for, as the device does not store that information. Non-resident keys replace passwords, but not usernames.

😈 Remember how I said I think we've "settled" on a definition? As far as I can tell the common parlance for "passkeys" is now "Resident" or "Discoverable" keys (indeed, passkeys.dev uses this definition). This is a problem, for a few reasons, a key one being how they interact with hardware tokens like Yubikeys. I'll talk about that shortly.

Okay, so now that is out of the way, let's talk about how a Relying Party (service provider / website / etc.) can implement passkey authentication. We'll start with non-resident keys.

The steps to do this are basically as follows (and this will look very familiar to the 2FA post because it's exactly the same process). For code examples see Part 2 of the series.

For Registration:

  1. The user goes to a website and clicks on "register", providing a Username in the process.
  2. The site responds by requesting the public key from the device using navigator.credentials.create (providing a few values from the server).
  3. The authenticator generates a seed, which is sent back in the PublicKeyCredential.id field of the response. The site must store that ID alongside the username and will need to send it back when logging in. It also generates a public key which will be used for signature verification later. Store that as well.
    1. Make this a one-to-many relationship on your user record please.
    2. If you're curious how YubiKeys any program can query every key on the device. You do not need to know the relying party. There's an operation to just get it to dump the service list. Generally this requires a pin or a biometric to unlock do this Duo wrote up a blog post
  4. The server must verify the credential using the 19 point process outlined in the spec (just use a library).
  5. You now have a non-resident key stored for the user.

In snarky diagram format:

sequenceDiagram
    Participant Passkey Device
	Participant User
	Participant Browser
	Participant Server
	User->>+Browser: I would like to Register.
	Browser->>+User: Username please!
	User->>+Browser: It's cthos
	Browser->>+Server: New user coming in, please give me an id and a challenge key
	Server->>+Browser: Here you go. "e98fowfi" and "bacon".
	Browser->>+Passkey Device: Please give me a new key for user ID "e98fowfi" and sign this challenge: "bacon".
	Passkey Device->>+Browser: Okay, here's a `PublicKeyCredential` with the key and signature.
	Browser->>+Server: Here's the public key, Id "1d889" and stuff
	Server->>+Browser: It all checks out, let them in.

To Log the user in:

  1. The user goes to the website, clicks "login" and then provides their username.
  2. The server retrieves the key IDs that have been registered with the account (this could be multiple keys). Then, it uses navigator.credentials.get passing those key IDs along in the allowCredentials stanza.
  3. Using the credentialId list, the authenticator will loop through and attempt to re-create the private key from the input. If it can do so, it will sign the request with the private key and send back a PublicKeyCredential like in the registration step.
  4. The server must validate that response and that the signature matches what it was expecting.
sequenceDiagram
	Participant Passkey Device
	Participant User
	Participant Browser
	Participant Server
	User->>+Browser: I would like to Log in.
	Browser->>+User: Username please!
	User->>+Browser: It's cthos
	Browser->>+Server: Please give me the keys and challenge phrase for "cthos"
	Server->>+Browser: Here you go. Key ID "1d889" and "justicier".
	Browser->>+Passkey Device: Please sign this challenge: "justicier" using key "1d889".
	Passkey Device->>+Browser: Okay, here's a `PublicKeyCredential` with that signature.
	Browser->>Server: Here's the response, is it valid?
	Server->>Browser: It all checks out, let them in.

And that's it! For a given user ID, something like a YubiKey will only generate a single key pair, but you can use the same key for different users with a different generated key. Password managers and Apple/Google can operate in this mode in much the same way (though I believe they generate and store a key).

So what's an advantage of using key derivation over storing keys each time? Well, something like a YubiKey only has limited storage space. In fact, most resident-key-capable YubiKeys can only store 25 of them. Likewise, by virtue of not storing keys on the device, someone who steals your YubiKey can't figure out what sites you use by probing the key.

Right, so how are resident (discoverable) keys different?

These follow a very similar process, but instead of needing to provide a username, you can instead query the device for valid identities.

To Register:

  1. The user goes to a website and clicks on "register".
  2. The site responds by requesting the public key from the device using navigator.credentials.create. This time you pass requireResidentKey as true (if you want to force resident key creation).
  3. The authenticator generates a key pair, which is stored on the device associated with the Relying Party ID you sent in the request and the user.id you provided. Like before, you get a PublicKeyCredential response that includes an id and the public key.
    1. Store this along with your user record and whatever user.id you send along with the request, you'll need it later.
  4. The server must verify the credential using the 19 point process outlined in the spec (just use a library).
sequenceDiagram
Participant Passkey Device
	Participant User
	Participant Browser
	Participant Server
	User->>+Browser: I would like to Register.
	Browser->>+Server: New user coming in, please generate a new id and a challenge key
	Server->>+Browser: Here you go. "99209ed" and "bacon".
	Browser->>+Passkey Device: Please give me a new key for user ID "99209ed" and sign this challenge: "bacon".
	Passkey Device->>+Browser: Okay, here's a `PublicKeyCredential` with the key and signature. I've stored it under `99209ed` and your domain name.
	Browser->>Server: Here's the public key for user `99209ed`
	Server->>Browser: It all checks out, let them in.

To Authenticate:

  1. The user goes to a website and clicks on "Authenticate with a Passkey or something" (there's also an autofill API you can use).
  2. The server creates its secret, like before, but this time when calling navigator.credentials.get leave allowCredentials as an empty array, or leave it out entirely.
  3. The device will query itself based on the relying party ID and provide a list of credentials and the browser will prompt the user which one to use (if there are multiple).
  4. The device will send back a PublicKeyCredential which will include the key ID, an attestation, and the user.id it was provided during registration.
  5. On the server, query for that id, pull the public key associated with it, and validate the signature using the 19 point process.
  6. At this point you now know the user and their info! You're here. Hooray.
sequenceDiagram
Participant Passkey Device
	Participant User
	Participant Browser
	Participant Server
	User->>+Browser: I would like to Log in.
	Browser->>+Server: User login coming in hot, please give me a challenge.
	Server->>+Browser: Here you go: "lemonaide".
	Browser->>+Passkey Device: Please sign this challenge: "lemonaide" using any key you possess for my website.
	Passkey Device->>+Browser: Okay, here's a `PublicKeyCredential` with that signature, and the user.id `99209ed`.
	Browser->>Server: Here's the response, is it valid for user.id `99209ed`
	Server->>Browser: It all checks out, let them in.

So, you've gained the ability to authenticate the user without them needing to remember anything besides their device.

This is good, generally, but it's also bad when you consider security keys. Because of their limited memory if every service were to move to the resident key model you'd either need to carry an entire key ring of YubiKeys...or we're going to need a bigger storage YubiKey.

Likewise, because the keys are stored you can list them. Fortunately this usually requires you to enter a pin or use your biometrics to do so.

😈 Here's the point where I rant a little bit. I prefer non-resident keys because I prefer physical security keys. If we get a physical key with unlimited storage that doesn't cost a small fortune, I might change my tune.

Like the MFA post, I want to recommend a Google Developers post, which will walk you through how to do this in depth, from both the server and client side.

There are a lot of sharp edges, but I'm confident you too can implement passkeys.

Looks like this did manage to be a fair amount shorter than the previous posts in the series. Nice.

For the final post in the series, Part 5, I'm going to cover some oddball scenarios and other things that don't fit neatly into the rest of the posts like machine authentication and non-web identity systems.

See you all next time!

Auth Series 5: The Other Stuff

Header showing the Kerberos and JAMF logos

Welcome to the last post in my end-of-2024 Auth series. This post covers the "other" stuff that I didn't cover in the previous 4 posts. It's mostly just the fun wrap up post to hit some things you probably won't need to know about but might want to know more about.

So, here's the list:

  • OIDC/OAuth other flows (M2M and Device Pin)
  • Certificate/Ticket based SSO, like Kerberos
  • Mobile device management and certificates

Let's get right in!

Remote control facing a blurry TV
CeltStudio @ Shutterstock #2179196459

If you have a "smart" device without an easy way to type in your password, like an internet-connected TV, or refrigerator, or light bulb (or something) you might have seen the OIDC device flow.

Basically, when you request to log into a service, you give it your username or email, and then it will show you a short (usually) code. You're expected to go grab your phone or computer, log into the service over there, and enter the code to bind your device to your account.

Auth0 has a nice flow diagram of this flow. The things you want to note about this are:

  1. The URL you use to generate the code is different from the usual URL you'd redirect users to in interactive mode.
  2. Your device has to poll for access to /oauth/token, there's nothing in the standard that gives a push (but I suppose you could implement your own push solution).
  3. How you bind a device to the account is up to you, and will depend on the device in question. Your application could generate it itself and store that locally, or you could use a unique device ID from the platform as appropriate.

OAuth 2.0 Resource Owner Password Grant (ROPG)

Permalink to “OAuth 2.0 Resource Owner Password Grant (ROPG)”

Did you know there's an OAuth 2.0 grant where you collect the user's username and password directly and send it to the resource server yourself? Well there is! You probably shouldn't be using this grant for anything, but sometimes there's no other choice. It's pretty much only okay in a first party scenario where you control both ends, but I've seen it used as a shortcut in a lot of places rather than showing an OIDC style redirect.

😈 The number of arguments I've gotten into about "just let them enter their passwords in the app, the popup is ugly" have been notable.

The flow is really short:

  1. Prompt the user for their username and password.
  2. Send that directly to the oauth/token endpoint and get an access token back.

Now, it's worth noting that this flow has been removed in OAuth 2.1 because it's not nearly as secure as the other options.

The last one I want to touch on a bit more than I did in the SSO post is the client_credentials grant, which is designed for use in Machine-to-Machine scenarios. It exists outside the context of a user, so this grant type tends to either be very privileged or limited to a small number of scopes.

This requires the pre-sharing of secrets between the authorization server and the client and as such it requires a client_secret. Ideally you get this secret to the client securely because anyone in possession of it will be able to log in with this grant.

The flow is dead simple:

  1. Send the client_id and client_secret to the /oauth/token endpoint and get an access token back.

How long that token is good for is typically configurable at the authorization server. Auth0, for example, defaults this to 8 hours (and they charge based on how many times you do this).

You'd use this grant if you have some backend that needs to interact with protected resources, and the client can be restricted to requesting certain scopes.

Group of Wolves
David Dirga @ Shutterstock #358007024

😈 I've never used Kerberos professionally so this is going to be pretty light on details.

Like other things I've talked about, Kerberos uses cryptography to use its magic, and it can work with either symmetric or asymmetric keys.

For the symmetric flow, it works kinda like this:

  1. The user sits down to a computer and types in their username (or provides their identifier some other way).
  2. The Authentication server does a lookup to see if that client's registered, and if it is, it'll encrypt a session token with the user's password (or private key) and send that back to the client.
  3. The client enters their password which is used to try to decrypt that key.
  4. Assuming everything goes well that session token is used for any further communications with the server.

There a few different ways to do this, all of which I'm only familiar with on the end-user side of things, but you'll find this a lot in corporate environments, and you usually won't have to think about it or even realize it's happening. One example is being able to log into corporate Wi-Fi and have any internal applications already know who you are. I've seen this implemented using a CBA solution, where a certificate is pushed to the device in question and that certificate is then used to authenticate the device to the network.

CBAs require you to set up a Certificate Authority (CA) that can issue certificates and is trusted by the device (which you can force via the next section).

😈 This is also how HTTPS works, but the CA is trusted by...general consensus that the CA is safe.

Basically, the way that this works is each certificate can contain information about the client encoded in it, which is signed with a private key that only the CA holds, so you can use that certificate to validate who the client is and that their client certificate has not been tampered with.

How this works in practice is that when you're authenticating with something, you can send that certificate and have your backend verify the Certificate with the CA and then make authentication decisions based on that. But, to get this to work with websites magically... you have to do some interception shenanigans. For example, having the device use that certificate instead of ... regular SSL certificates.

And how might one do that? Let's talk about Mobile Device Management.

If you're working at a big company and your work computer was issued by an IT department, you probably have some sort of MDM installed on your device. MDMs allow a centralized server to do all sorts of things to that device, like install applications, remotely wipe the machine (in the case of loss or theft), or identify the machine to the network. We're going to focus on that last part.

One major MDM vendor is jamf, and it's the one I've got the most experience as an victim end-user of. On Mac, it uses the same certificate process to gain control of the device (and shows you exactly what it can do in that certificate).

😈 Your machine almost certainly also includes some sort of monitoring solution to look for policy violations. Assume anything you do on a work machine is being monitored.

So, remember the section right before this where you can install a trusted CA on a device? Using an MDM is a way you can do this (you can also image the machine before sending it to the end user, but then you have to worry about updates somehow). So, here's how an MDM facilitates the magical auth from the section above.

  1. Have the MDM solution identify a user via whatever auth mechanism you want (installing an MDM profile and then having them log in for credentials is what I've seen).
  2. The MDM and the CA create a certificate for the user and device and pushes that certificate to the device.
  3. The MDM configures the client machine to both trust and use that CA for all of its traffic.
  4. Additional shenanigans are in place to configure say a browser to use that certificate essentially as a man-in-the-middle to intercept traffic and present that certificate to a website or internal application to do auto-logins.

That does mean that the traffic going to those (or all) websites is encrypted by the custom certificate meaning the company can see encrypted traffic to those websites. For an example of someone figuring this out, check out this Stack Exchange post. Because your certificate is also unique to you, they can introspect that traffic.

😈 BTW big companies are probably doing telemetry on this traffic, but in my experience they're not going to proactively go poking around in the giant piles of data unless they have a reason to.

But yeah, it is pretty magical not having to type in your credentials to business applications.

Right, we're done here. I hope you enjoyed this series. If you'd like to see more long-form deeper dives on topics, let me know on the socials or here in the comments.

Mozilla is now an AI company


Shutterstock #1956259915

In an ongoing series of my disappointment with this AI bubble, Mozilla has gone all-in on the AI Hype Train. Their thesis is dodgy, it's basically the "if we don't build AI into our browser, the evil big corporations will be left unchecked". Right. Okay. It's worth noting that this extends on their idea of "Responsible AI" from a white paper they published in 2020.

😈 If you're looking for a browser that doesn't have any AI features in it, I use Vivaldi. They've committed to not shoving AI into everything. Also consider donating to Servo, an alternative rendering engine.

There's a fair amount to unpack here, but I want to start with a quote from their 2025 updated white paper which I think is telling (emphasis mine):

Consumers are starting to pay attention to AI’s impact on their lives. Workers — from delivery drivers to Hollywood writers — are pushing back on how AI affects their livelihoods. However, we have not yet seen a wave of mainstream consumer products that give people real choice over how they interact with AI. This is a key gap in the market.

There is a big presupposition built into this: Generative AI (that's the thing upending livelihoods) is a technology that deserves Mozilla's attention, and could be a market differentiator for them.

How? Like, really, how? I do not believe that there exists, nor can exist, an "ethical LLM". The current conversation is how LLMs are driving people into psychosis and encouraging vulnerable people to suicide, as if that's just because the companies making the LLMs just don't care enough. I'd argue that using a very fancy next-token-generator runs that risk regardless of who's in charge of it. Sure, OpenAI and others have nudged the LLM into seeming deferential and conciliatory, but that's part of why they are so addictive. It's very unclear to me why Mozilla thinks they can make an LLM that doesn't do these things that also provides the illusion of competence.

Based on their stated goals, you might think that Mozilla has some exciting trick up their sleeves to make LLMs safer. You'd be wrong. Their addition to Firefox so far has been:

Taken from the Mozilla's page on LLM Chatbots

Look at all those big tech LLM models the browser prompts you to try out.

Perplexity search engine

I'm sure Perplexity is "ethical" and this integration is surely going to help people who are already using it.

Tab groups gif

Great, cool, let's spend the compute calling an LLM to auto name tab groups.

Apparently we needed this

Because Open Graph tags are too much to ask.

Seriously, who? If we take an argument I saw on Bluesky, "Millions of Firefox users use ChatGPT and thus they should have an option in Firefox"[5]. Citation really needed there, because near as I can tell most of Firefox's users (as an alternative browser) by-in-large do not want AI features. As of October, 2025 - Firefox represents about 4% of browser market share. According to some website statistics[6], chatgpt.com gets about 2 billion visits a month. If we assume that traffic is evenly distributed among browsers (it's not) and that every visit is a unique Firefox user, you'd get about 80 Million Firefox users using ChatGPT! Woweeeeee. But it can't be that high, those are not unique visits. We have no idea how many actual unique visits they're getting, and that'd be 10% of OpenAI's reported number of "users".

All of that, however, is beside the point - going to a website does not equate to "want to have that website deeply embedded in my web browser". If that were the case, we should have a dedicated PornHub sidebar in every browser.

The other argument I've seen is that "well, no one is asking for it, but we don't just design features based on what people ask for". This is basically the same argument as "if you asked someone in the past, they'd have said 'build a better horse'". Which has always struck me as funny. Car usage was a choice, and had we not invented the Automobile we might have had much more accessible cities (because there wouldn't have been a concerted effort to defer to car companies). Who's to say a world where we've got faster horses is better or worse than the one we currently live in?

Now, there are some boosters out there who make the argument that "people are going to use LLMs anyway, so we should have an 'ethical' company providing them an alternative." It's funny that this "alternative" always includes "more LLMs" or "more AI". Throw more technology at the problem, and that is the solution. The answer is apparently not "provide them a way to opt out of Generative AI".

No, to the SV tech community, the thought that the answer could be "Provide alternatives to the problem that are not LLMs or diffusion models" is anathema. The answer to the "busy mother looking for recipes" is a search that hasn't been inundated with slop results, that provides links to sites in context.

The answer to "people are developing a psychosis when talking to a chat bot" is not "MAKE IT EASIER FOR THEM TO TALK TO WHATEVER CHAT BOT THEY WANT!". It's developing universal healthcare and easy access to trained professionals. Not giving them an ad-lib machine that talks like people but only generates a stochastic output.

😈 This is _exactly_ what Mozilla has done so far. With all their talk of addressing the "problems with model training" their foray into the realm is literally "add a chat sidebar".

This is a position that implies inevitability

Permalink to “This is a position that implies inevitability”

A certain discourse has taken hold: "I'm not saying AI is inevitable" while suggesting that the way out is more LLMs. I think what these people are saying is really just "Corporate controlled AI is not inevitable, but LLMs are." The very idea that you have to incorporate an LLM into everything because it is "popular" is an inevitability argument.

I don't see why they don't understand that. Are they so surrounded by people who think LLMs are a path to full societal automation that the idea that they're a dead end being propped up by a ridiculous amount of investor money that they think the rest of us are idiots?

Perhaps (as Ethan Marcotte suggests) LLMs are a failed technology and shoving them into everything to keep a bubble inflated is a terrible position to hold?

Maybe these folks think that technology is neutral, and LLMs are only as bad as the people that are pushing them, and not that they are The New Aesthetics of Fascism?

So what are you doing, huh? Laws won't work

Permalink to “So what are you doing, huh? Laws won't work”

Listen, LLMs and Diffusion models are causing a lot of direct problems right now, and if your solution to those problems is "But what if I made another LLM", then we are never going to agree. I do not believe you can create an LLM that is built on ethical consumption and have it look anything like what we've come up with.

I feel like the crowd that says "well saying they suck hasn't fixed anything" is being pretty cynical. They don't think there's legislative appetite to do anything about it, and they might be right - for the moment - but there is a deep appetite for political change right now, and more progressive and, dare I say, socialist candidates. If we can salvage what's left of our democracy we could, in fact, do something about this collectively.

Instead, their solution is just another techno-feudalistic idea wrapped in "but what if this company doesn't turn out evil".

I'm just so tired of that line of reasoning.

I honestly don't know. They've been making a bunch of decisions I don't agree with for years now and I honestly have no clue other than "a lot of execs think that LLMs are a path to a machine god and thus think it needs to exist". The folks over at DAIR have written extensively about the underlying ideologies fueling this bubble and... I think it would behoove Mozilla to read up on just what their goals are, because proliferating LLMs (and Diffusion models) only helps them.

Anyhow, yeah, I'm tired. So tired.

For anyone who feels like I do, keep it up. There aren't as few of us as the boosters think, and the bubble is running on the fumes of lying about capabilities.


  1. https://support.mozilla.org/en-US/kb/ai-chatbot ↩︎

  2. https://www.reddit.com/r/firefox/comments/1lbu91u/its_official_mozilla_quietly_tests_perplexity_ai/ ↩︎

  3. https://support.mozilla.org/en-US/kb/how-use-ai-enhanced-tab-groups ↩︎

  4. https://support.mozilla.org/en-US/kb/use-link-previews-firefox ↩︎

  5. This is a paraphrase of a couple of different arguments smashed together. It doesn't represent the natural and complete thoughts of any given individual, and it annoys me that I have to say this. ↩︎

  6. https://www.semrush.com/website/chatgpt.com/overview/ ↩︎

LLMs are a Cognitohazard

Cognitohazard image

So, I've been watching a bunch of folks whom otherwise have reasonable takes about LLMs who, under very specific circumstances, have extremely weird takes about how much an LLM is helping them do a thing, and their internal justifications for them. This isn't about anyone specific; for each of these points, I've seen variations of these from more than one person in every case. But they're happening enough that I want to talk about them. Don't just take it from me, this post from @xgranade makes the point better than I can:

It is amazing how many people I see claiming that AI has one good use. No one seems to agree on what that one good use is, and there's always a hell of a lot of goalpost shifting and special pleading involved.
For some reason a lot of folks quite reasonably follow the arguments against AI, but then partition off the one thing as immune or exempt from having to worry about any of the ethical and practical problems. - @xgranade@wandering.shop

Let me start, though. I think each and every one of these things ignores the externalities of running an LLM: The wholesale scraping of data frequently without consent (and the associated pushing of costs onto the site being scraped), the massive build-out of data centers that will never be used, a global chip shortage because the AI industry can just do that, the proliferation of slop and disinformation, and the trend of people experiencing psychosis when interacting with chatbots.

Okay, with that out of the way, let's dig in.

These are paraphrased. They are not direct quotes. If I use quotes, they are scare quotes to separate them from the rest of the sentence.

LLMs are great for Prototypes - I'd never have the time or ability to do this. It'd take me forever, or I'd never get it done.

Permalink to “LLMs are great for Prototypes - I'd never have the time or ability to do this. It'd take me forever, or I'd never get it done.”

Okay, so this one I've seen a whole lot lately. Namely, the claim that busy engineers are able to rapidly prototype something and get real time feedback on the prototype. This is usually associated with a sense that the engineer in question is "not a {insert type of} engineer" and so they'd never be able to complete such a task.

There are a bunch of problems with this approach, but I want to start with the claim that these folks would never be able to do something. That's patently untrue. If you're a software developer you absolutely have the capability to learn how something works and build out a prototype. The LLM's statistically average output is going to produce code that is prevalent. The most prevalent front end code being fed through these things is React. React has some very silly contrivances in it, but it's not an insurmountable mountain to learn the basics of React to roll out a prototype.

For code that is more difficult to learn, an LLM is going to have a harder time than you would, because the amount of samples in its training set are probably going to be a lot smaller.

The second problem with this attitude is that if you acknowledge that the code an LLM generates is substandard, and you want to treat it as throwaway.... it is never throwaway. Prototypes tend to make it, unmodified, into production. And if your usual engineers are busy, no one is going to substantively review this code.

Another issue is that the human involved in this process is not going to learn much from this process. Having the LLM write the code, and then try to make sense of the generated code is a failing proposition. The LLM does not have cognition. It produces code that is statistically likely. That's going to be a twisted mirror of the code it was trained on. Will it execute? Maybe! It might even produce code that executes on the first go. That's how slot machines work.

Also, this is why scaffolds and quickstarts exist?

I have the LLM do initial code review for me and point out things that are wrong

Permalink to “I have the LLM do initial code review for me and point out things that are wrong”

This one is the least offensive to me on this list, but it's still super annoying that I have to, for example, review a bunch of suggestions that Copilot makes on a GitHub PR that are either subtly wrong, or correct enough that accepting its suggestion won't break anything but it's also not doing anything of value (for example, making a comment more verbose).

I'm not convinced that this actually saves more time than it wastes. I've not had it find a single bug that would have caused a major problem had it made it into production, and it has (had its suggestions made it through) introduced bugs that would cause problems.

Perhaps I'm also being a bit petty, I can discuss a problem with a human and get to the point where we understand where there's a misunderstanding on either side - the LLM doesn't do that, it'll just agree with whatever you type to it, whether or not you're actually wrong. Because it's statistics.

It's worth pointing out that you, writing good unit tests and utilizing static code analysis (linters, anyone?) should be able to catch the same amount of stuff that an LLM would, deterministically.

Oh gods. Okay. So. I'm also a writer (fairly novice, granted), and I've found that there's nothing about typing at an LLM that actually generates quality ideas. It just spits out mild rehashes of old ideas and frequently the same variant of an idea a bunch of times in a row.

Why would I use electricity for this? I can grab a coffee, stare out the window, or talk to an inanimate object on my desk and generate a ton of stupid ideas that I can then iterate on, and I've found that works out way better than running a GPU to spit out "ideas".

I encourage you to give this a try. Find somewhere quiet, stare out of the window, and let your mind wander. Ideas will come.

LLMs are awesome at summarizing! I'm a busy professional and there's a lot of stuff for me to read!

Permalink to “LLMs are awesome at summarizing! I'm a busy professional and there's a lot of stuff for me to read!”

Nope. LLMs do not summarize, they shorten. They also frequently introduce inaccuracies or make up quotes when they're doing it. Often it'll also remove critical information, or emphasize the wrong thing. Overall, you cannot trust even the most advanced LLMs to accurately summarize a longer piece of text for you.

So what do you have to do if you want to be sure you have an accurate understanding of a longer text? You have to read it yourself. Which defeats the purpose of having the LLM shorten the text for you in the first place.

If you don't care about the accuracy of the summary...why are you reading it?

LLMs are great as a generalized search engine!

Permalink to “LLMs are great as a generalized search engine!”

So, this one is weird because it kinda depends on what you mean by "search engine". An LLM by itself isn't a very good search engine. It's maybe an okay reverse thesaurus or dictionary, but if you're looking for something with up-to-date information you're probably taking about something like Perplexity or the like.

Note, you can ask an LLM to generate "citations" and "link to your source" but it's reasonably likely to hallucinate those.

In the case of Perplexity and Google LLM search, what's actually happening behind the hood is a bit...insidious. They're using Retrieval Augmented Generation, namely, doing a web search first, and then feeding the first couple of results into the LLM to have it "summarize" (see previous comment) the results, and using a couple of techniques to inject the source reference into the "summary". Couple of major problems with this:

  1. It's only going to use the first couple of results for its summary, lest it risk overrunning its context window.
  2. It removes the results from their original context, so you can't be sure the summary doesn't include a shitpost from Reddit.
  3. "Hallucinations" still can crop up in those summaries.

Is this better than Google search has become? Maybe! Google has gotten pretty dang bad. I've resorted to running SearXNG on my own infra to get less bad results and no summaries.

Relatedly, we're seeing more and more hallucinated citations appearing in places they shouldn't so, uh.

This is just placeholder dialog, it'll be removed before the final product is out

Permalink to “This is just placeholder dialog, it'll be removed before the final product is out”

This one is a two-parter. Namely, the using of an LLM for "filler" text, and the use of a text-to-speech model to have verbal dialog - most commonly in games.

The first is a problem for even if you intend to remove it, bogus "filler" text will blend into the background. It will leave its mark on the final product. The odds that you'll remember to remove it decrease, when you could just leave the prompt in ("Emotional scene where Santa admits he made a pact with a devil for immortality", e.g.) and it'd be a lot clearer what needs to be rewritten. By not doing that, you're creating background radiation. Mediocre stuff that you may just miss when you're doing your next dialog pass.

Like a lot of the use cases I see for generative AI, it boils down to "just send me your prompt" rather than having the token spewer spit out tokens.

The latter is just... Like. If you need filler audio so you can "hear" the scene, just have someone around you read it out loud? Record it yourself? The same thing as before applies, if you don't intend to keep it in, there are other ways to get that out that do not require much more effort. You'll get more emotional nuance from a person who understands the scene, rather than the flat garbage that Amazon tried to pull.


That's all I've got energy for at the moment, but there are definitely more of these floating around out there, so I'll probably update this post later to add more, or make a follow-up. Haven't decided. I'm just so tired.

Asahi on Macbook Air M2

Asahi and Macbook
Shutterstock #2041259501 and Asahi Logo

Over on Mastodon I mentioned, offhand, that I'd been daily driving (meaning, using daily but not as my primary machine) Asahi on a Macbook Air M2 (15") and some folks were curious what kind of experience I was getting with it. So, instead of making a giant thread, I figured I should blog about it instead given just how behind I am on blogging. Plus, most of my blog posts are lamenting the state of slop so this should be a nice change of pace.

I'll just walk you through what I'm using it for and the process from installing Asahi to using it on a daily basis.

Right, so this is worth noting that this is not my primary desktop. That is either a Mac Mini M4 or a System76 Thelio depending on if I'm at home or in my office. I use both of those for primary work, including destkop publishing stuff (which we'll touch on shortly because it's relevant to Asahi).

The laptop is used as a "writing machine", which means it has to do the following:

  • Browse the internet
  • Run Obsidian
  • Run Nextcloud Desktop (to sync my writing files)

And that's it. However, it's worth noting that I also have git and some light dev tools installed on this so I can do stuff like update this very blog from this machine. For my editor I'm using Lazyvim with a bunch of plugins these days, and that runs basically anywhere.

Notably I'm not trying to hook this up to an external monitor (nor have I tried), and last I checked this was still a bit hit-or-miss if you're expecting displayport alt mode to work. Even if it does, it only supports a single external monitor due to hardware limitations.

So, if you're expecting that to work well, this is likely not going to be viable for you.

Installing Asahi was a breeze. The Asahi team has a shell script that goes and downloads a secondary package from their CDN and then runs an install shell script from there. The source for the installer is on GitHub if you want to go poke around in there and/or download it directly. The shell script from alx.sh is just a thin wrapper around downloading and executing said installer.

Once you're running the installer it interactively walks you through what you need to do:

  • Resize the disk to make space for Asahi
  • Install Asahi
  • Reboot into recovery to finish the installation.

There's a Youtube Video showing the full process, if you're interested in what it all looks like.

WARNING: You really need to pay attention to what it is telling you. You have to take a manual step after the install is completed in order to get this working. It's recoverable if you don't, but it's a hassle and can render your mac unable to boot without reinstalling MacOS, but pay attention.

After that's all done (especially the manual recovery step to get uboot functional) you'll be taken into the Asahi Fedora Remix setup flow where you set up the user account and all the regular Fedora stuff.

Okay! Once that was done, the next thing you want to do is install some apps. Discover is shipped by default, and if you can find the app you're wanting in there, you're golden. Flatpaks included.

dnf is also available if you want to install packages that way.

So, I wanted to get a few things running immediately:

  • Vivaldi
  • Nextcloud Desktop
  • Obsidian

The only problem with that is that neither Vivaldi nor Obsidian are in Discover because it doesn't have Flathub enabled by default. To do that, you can just go into discover's settings menu and click "Add Flathub", and viola.

But because I'm a masochist, apparently, I figured "You know what, let's just go get the AppImages for these two things I'm missing, like a fool, and run those!" The thought apparently never occurred to me that I could also just download the Flatpakref directly. I'm not sure what I was thinking. Probably something like "ooh I wonder if the AppImage will run okay." They do.

Dear reader, that does work, but once you've done that you want to move them somewhere safe and then set up menu entries for this. Normally I use Gear Lever for this, but since I chose not to add Flathub, that wasn't an option. Good news though, KDE offers a handy menu editor so you can just add whatever you want "easily" and so, like a fool, I did that.

Nota Bene: I've since tested the Obsidian flatpak on Asahi. Works fine.

The other bit of software that I use, Freetube, was giving me some trouble out of the box, and that required me to go in and add a command line flag to the menu entry: --js-flags=--nodecommit_pooled_pages which I pulled out of this GitHub issue which happens on other Electron apps as well, which is apparently to do with Asahi's page size.

Fortunately, I've not encountered anything else that's super broken when installing stuff, it largely "just works".

The thing I wish I had gotten working but haven't been able to solve for

Permalink to “The thing I wish I had gotten working but haven't been able to solve for”

Okay, one thing that I would love to get working but cannot figure out is getting the v2 Affinity suite running on one of these things. It's about the only reason I have to boot back into MacOS to do.

Getting Affinity running on Linux is a bit of a chore, but thanks to AffinityOnLinux it's very straightforward and I've done it on several Linux installs (remember that System76 desktop from earlier, yep, it's working there).

Unfortunately, that's not so true on Asahi, where not only do you have to contend with Wine, you also have to translate from x86 to aarch64. Now, this is theoretically doable, but any combination of things I've tried has not worked.

Alas. Perhaps I'll give it another go later.

I can't really complain about the battery, it seems to last almost as long in active use as MacOS does.

What it does not do is sleep properly. If you just close the laptop lid, expect the battery to drain just as fast as if you were actively using it (maybe a bit slower since the screen is not on). You cannot just close the lid and expect it to still have 95% charge after a week of it sitting unused.

So, remember to turn the thing all the way off.

Asahi does this fun thing with the cmd key where it'll work in a mac-like fashion for some things, but not for everything so it really mucks with the muscle memory. This isn't a huge deal for me, because I frequently switch between Mac and various flavors of Linux, but I can see it messing with someone who's just hopping over for the first time.

Haven't had any issues with Bluetooth connecting to my earbuds, nor have I seen any weirdness with the built in speakers or adjusting the display brightness or keyboard backlight.

Resuming from the screen turning off does occasionally require me to tap the power button, but otherwise there's noting of note there.

Touch ID does not work, of course, there's no support for the fingerprint sensor here. Knew that going in, though, so it's not a downside.

I run Tailscale on it for Wireguard stuff and that works great. Wifi's been solid.

I'm sure if I had some more demands of Asahi I'd run into some things that are annoying, but for what I'm using it for, it's been really nice and generally aligns with my overall expectations of running Linux on a modern Macbook.

As an aside, I also have a 2015 Macbook Air 11" that runs Arch... it's a neat little machine.

So overall I think it's worth giving a try if you don't have really high demands of the thing, and even more so if it's not your main machine. I could probably get by with just Asahi for the kind of computing I do for my freelance work, but I've never tried it for that partly because of dealing with various VPN things.

... and yeah, that's it. Lemme know if there's anything you'd like me to answer on this post and I can update with an FAQ later.

I've got a Novella Out Now

Weight of a torch banner image. Robed figures holding real big candles.

The title pretty much says it all! I've got a Novella (really, it's five short stories in a novella format that are directly related to one another) out now at The Arcanist: The Weight of a Torch.

Three of the stories are available on the Arcanist's website if you'd like to get a feel for the Dark Fantasy vibe.

If it's your jam I'd love honest reviews on Goodreads or the Amazon page.

The Weight of a torch Novel Cover