3982 stories
·
4 followers

The search for a Spotify alternative

1 Share

About six months ago, I realised I was feeling increasingly uncomfortable paying for our Spotify Premium Family plan. The company had hiked its prices yet again. It had changed its policy to remove any royalty payments to artists with under 1,000 streams a year — despite already taking the number one spot for the worst-paying streaming service (already a low bar). And Spotify CEO Daniel Ek is investing the huge personal profits he’s made from these tactics into AI-powered weaponry. So I decided, like quite a few folks have recently, that it was time to look for an alternative streaming service.

And in those six months, I’ve been trying pretty much all of them. This post is intended to serve as a non-exhaustive list of considerations for anyone else of doing the same. (Because it was exhausting, believe me.)

But first, a caveat: this post is about choosing a streaming service, not about if they could or should be used in the first place. I’m a lifetime Plex user and I love it. My Plex library is loaded with all the music that’s not on the streaming platforms, most of it purchased from Bandcamp. But, for better or worse, streaming is a must-have in our household (and switching or remaining on a service is a four-person decision), so going Plex-only isn’t an option.

Second caveat: as an artist, I have no intention to remove my music from Spotify. Although I’m not so active with releasing new music these days, taking my music off the platform would just be shooting myself in the promotional foot. So that’s a consideration for another day.

Lastly, a warning: please don’t do what I’ve done, which is pay for multiple services for six months while I attempt to make my mind up and write this bloody long blog post. Oh, and here’s another tip: TuneMyMusic is what I used and (temporarily) paid for to move my library to and from the various services — although it’s worth Qobuz has Soundiiz integration built in for free.

Oh, and: all of these views are my own, obviously.

The logical first choice in a search for a Spotify alternative was Apple Music. We’re already paying monthly for Apple TV+ and extra iCloud storage, so upgrading to an Apple One family plan would only be a couple quid more a month. Easy decision. I’d tried out the service about a year or so ago and not got on with its design, but guessed that in that time, things must’ve improved. Surely they would’ve finally found a way to get around that legacy confusion between Apple Music, iTunes, and the iTunes Music Store, right…?

Conceptually, it feels like the service is still unsure about what it wants to be, and this manifests in a variety of UX inconsistencies.

Try searching for an artist: you can filter your results by ‘Your Library’, ‘Apple Music’, and ‘iTunes Store’. That seems logical enough, except that Your Library is actually a mix of music you’ve either purchased or loaded into the-library-formerly-known-as-iTunes, plus any music you’ve saved from Apple Music.

Screenshot of the Apple Music service Left: an artist page within your library. Right: an artist page within Music.

Navigate through your library to your artist of choice, and suddenly it’s gone all iTunes-like again, where no artist names are links… unless, that is, you go via the Albums tab in the sidebar, and then they are links from those album pages… which now look somewhat more Apple Music-like again.

Screenshot of the Apple Music service Left: an album within your library. Right: the same album on Music.

It’s a frustrating experience because this inconsistency leaves almost every interaction feeling like a guess. And it’s even more infuriating because it feels like Apple should’ve fixed this stuff years ago. Not just fixed — nailed it. It’s Apple. They can and should operate the very best music streaming / purchasing / collecting experience there is, especially on their own platforms. Even with the superior design of the iOS app over its macOS counterpart, the conceptual heart of Apple Music is, in my opinion, still too muddled to be usable. Most of my friends who’ve moved there from Spotify seem like they’re putting up with it rather than enjoying it. (One thing I do like about Apple Music, which more services should support, is making the record label a link to all releases from that label.)

In fairness to Apple, one of my additional strikes against the service was because personally I just like dark interfaces for music apps. I’m sure this is because Spotify has accustomed me to it over the years, but the fact that Apple Music in black can only exist if you change your whole system to Dark Mode feels like an unnecessary restriction. So, with a dark interface as a criteria, my attention to turned to… well, pretty much any other streaming service.

Deezer — great, but let down by one huge bug

Deezer wasn’t originally on my list because when I’d last looked, its design looked pretty dated, but hold on, what’s this? Deezer now looks, in fact, rather lovely! Who knew? (Okay, lots of people knew — their brand refresh happened in 2023 and was handled by none other than Koto, and the brand’s bespoke typeface, Deezer Sans, was designed by the wonderful NaN, so of course it’s gorgeous.) Anyway, yes, this won Deezer lots of points in my book. I’m a pathetic, shallow, aesthetics-focussed designer, after all.

First impressions were pretty good. Deezer’s got a friendly interface with lots of customisation options (made dark instantly, of course). It’s kind of odd that the playlists sidebar only shows some of my playlists, but I can live with it. Slightly weird that clicking on the track titles does absolutely nothing: unlike every single other service, which effectively treats the whole track area as a button, you need to explicitly hit the play (or pause) button next to the track. Why is this? This hit area decision can’t be intentional, can it? It’s fine, though. I can live with this, too.

There are some lovely touches, like being able to pin artists, releases, or playlists to the ‘Quick access’ sidebar. The overall design and behaviour of the apps won me over: I decided Deezer was the one to switch to. Time to start the move!

But then something very odd happened.

I was on the profile page for Recondite, an electronic artist whose album Hinterland (released on Ghostly International in 2013) is one of my all-time favourites to play in winter time. But this album was nowhere to be found.

Screenshot of the Deezer music service The album isn’t there. In fact, no albums are there! This artist has albums, trust me.

Worried that it might not be in Deezer’s catalogue, I searched for it and found it. Phew!

Screenshot of the Deezer music service Wait, it is there!

But why wasn’t it appearing on his artist page? In fact, why weren’t any of his albums showing up on that page? After all, they do show up on the search results page for “Recondite”.

Screenshot of the Tidal music service I’m losing my mind here.

Going the other way, clicking the artist page link from the album page, demonstrated it was definitely connected to the correct artist, so I assumed it must be down to some inconsistent metadata. Oh well. Annoying, but not a biggie.

Except that then, it happened again. And again. And again.

These next examples were for more mainstream artists, and it started to seem pretty strange that the albums weren’t showing up on these artists’ pages. And it couldn’t be bad metadata — on every other service I tested, the releases were present on the profiles.

A quick bit of Kagi-ing revealed two things: firstly, that I wasn’t going crazy (what a relief); secondly, that this has been a known issue for years. Just take a look at the number of people questioning it on Reddit or even on Deezer’s own community pages (where several admins mistakenly claim that the issue is fixed). This came as a relief, but it’s also baffling: how can a service as mature as Deezer have such a fundamental flaw with its catalogue?

For folks questioning why this might be a problem, it’s worth remembering that exploring an artist’s discography is a vital part of the process of discovery — you find a new band you like and you want to see what else they’ve released — and therefore it’s a vital form of potential revenue for the artist. The money made from streaming is pitiful, whatever the service, but 10 streams of an album’s tracks that then convert to 100 streams of their other albums is clearly better than a one-release dead-end.

Tidal — nice, but not immune to metadata muddles

Deciding that this well-known and oddly unfixed discography bug made Deezer usable for me, I turned my attention to Tidal. Despite perhaps looking a little too much like Spotify, I immediately felt at home, and the inclusion of high-quality audio for a monthly price that’s a good 20% cheaper than Spotify’s seems quite reasonable. There are some nice unique features, too, like being able to group playlists into folders.

Screenshot of the Tidal music service Playlist folders! Why doesn’t everyone have this?

Unfortunately, Tidal isn’t immune to weird metadata bugs. I’ve come across a few instances of multiple profiles for an artist or band, with the releases spread across both, often with one very clearly being the official, label-managed page.

I noticed some of these being fixed. Even over the course of writing this post, Tidal seemed to consolidate Spiritbox’s profile to include all of their releases. But then I discovered Vower and found their discography split across two profiles.

Screenshot of the Tidal music service Left: Vower’s artist page. Right: also Vower’s artist page. Wait, what?

This actually highlights a problem with metadata as a whole: if you want, you can upload a release (via a distribution network), put in whatever metadata you like, and effectively break the system. This is exactly what’s happening with all the AI-authored slop being uploaded to the streaming services, credited to ‘real’ artists — and therefore receiving a tonne of streams from devoted fans keen to hear their new releases — when in fact the only thing shared by these nefarious tracks and the legit artists is the value in a cell in a metadata spreadsheet.

However, this is an issue with the way streaming services work rather than Tidal specifically. And Tidal does at least seem to be attempting to fix these things. So it looked like Tidal could be the one. Some slight weirdness, but no outright showstoppers like Deezer.

But then I got in the car.

Our Volvo XC40 has Android Auto. I can use CarPlay if I plug the phone in, but my wife and kids prefer something more instant, especially when you can just ask the car to play music. We’ve always used the Spotify app without issue, but for some reason both the Deezer and Tidal apps are more basic: searching for an artist and then tapping on that result simply plays their top track. Digging into any discography is almost impossible unless you ask for a specific release name. And the CarPlay apps aren’t actually much better because you’re forced into navigating via Siri — still the most useless assistant out there.

With frustrating apps getting in the way of a decent listening experience, maybe it was time to look again at some more alternatives?

YouTube Music & Amazon Music

YouTube Music: Around the time I started this experiment, YouTube had offered me a one-month free trial of YouTube Premium, which includes full access to YouTube Music. Or, to put it another way, your YouTube Music subscription will remove ads from your videos — and that’s a tempting offer. But I quickly decided against this service: the lack of a dedicated macOS app and a catalogue overstuffed with bootlegs made YouTube Music a no-go for me (even with its dark interface).

Amazon Music: Apparently — and I think this’ll come as a genuine surprise to just about everyone — Amazon pays artists better than some of the other streaming services. But even a very quick test of it proved that it was missing vital albums from my favourites. It was a consideration for about half an hour of testing.

Back to Apple Music via DaftMusic (and Albums)

Frustrated by Deezer and Tidal, and put off almost instantly by the alternatives, I was encouraged to see the release of DaftMusic — a UI for accessing Apple Music without any of the, well, Apple Music UI. And it’s nice! It manages to get around the iTunes-like ‘Collection’ weirdness and there are a load of customisation options, too. Plus, I love supporting indie developers.

Screenshot of the Daft Music app DaftMusic: a third-party UI for Apple Music. Way nicer than what Apple came up with themselves.

Unfortunately, the need to manually import playlists to the app (rather than them being synced from the Apple Music itself) is a pain when using multiple devices. I’m excited to hear that an iOS app is on its way, though, and how this might change things.

I also briefly tried Albums after my mate Jon Hicks’ recommendation. The iCloud-powered syncing across Mac and iOS is great and, again, yay for indie devs. But my family love their playlists, so an album-only interface is never going to fly.

Is anything as good as Spotify?

About two months ago, I was getting ready to hit ‘publish’ on an earlier draft of this post that ended with a disappointing conclusion: that, because every other streaming platform I’d tried has a range of issues — in some cases, bugs that hamper basic everyday use — it was impossible, right now, to replace Spotify.

Yes, it’s still susceptible to the metadata issues and abuses I’ve detailed above. Yes, it’s bloated and full of podcasts and audiobooks and videos you don’t want. And yes, Daniel Ek’s investment choices — but it does so much well. Custom-sorting for playlists feels like such an obvious feature and yet it’s missing from the competition. And Spotify Connect — being able to swap playback between devices seamlessly (something we do a lot in our house, especially when one iPad dies, or even to see when someone’s playing something in the car while I’m at my desk) is something I missed while trying every other service. But most importantly, it… just works. There are no weird bugs that stop me from discovering an artist’s discography or outright stop me from playing music. It’s far from perfect, but it comes a lot closer to perfect than the competition.

So I was about to end this experiment by cancelling all of the subscriptions I’ve been testing and decide, reluctantly, to continue paying for Spotify — a conclusion I wasn’t at all happy about. This prompted me to give one last service a try: Qobuz, the one billed for audiophiles and fans of Classical music. I’d dismissed it based on its marketing materials, but figured it was fair to give it a shot.

And I’m so glad I did.

The surprise twist: Qobuz is pretty damn good

I’m very happy to be publishing this blog post with a conclusion that feels morally right: I’ve cancelled our Spotify Premium Family plan and moved us over to the Family plan on Qobuz. It’s not perfect, but it does get most things right.

To be honest, it’s not the high fidelity sound (Qobuz’s main selling point) that really interests me; it’s more the thoughtful approach to the service as a whole. Plus, it’s customisable, and this is important because I must admit I’m not in love with its design out of the box. But modified to use the Qobuz Theme v1.3 by Jon Hicks, the desktop app is subtly but significantly better.

Screenshot of the Qobuz music service Left: the Qobuz macOS app, out-of-the-box. Right: the beta version of the app, using Jon Hicks’ Qobuz Theme v1.3.

Please note that the screenshots that follow all show this customised version, sporting Jon’s CSS. Is it fair to compare the other services to a version of Qobuz that isn’t actually representative of the company’s own vision for their product? Probably not, if I’m being honest. However, the fact that it is theme-able, and that you can get this UI with very little tinkering required, is one of the many benefits of Qobuz over the others.

Here are some things that stand out to me about this service over the others:

Qobuz Connect: Working exactly like Spotify Connect, you can move your music between devices with ease. It’s a little buggy in that one of my work Mac seems determined to return to the in-built speakers when I’ve paused music for a while, but it works well enough. I can’t understand how it’s only Spotify and Qobuz that have this output-swapping functionality — it should be on every single music service.

Screenshot of the Qobuz music service Qobuz Connect gives you a lot of choice when it comes to outputting the audio.

Releases: Qobuz calls them what they should always be called: releases. You can’t group all music as ‘albums’, as so many services do, despite them being EPs or singles or compilations. Okay, it’s a small point, but these things add up to show that they care about music.

Screenshot of the Qobuz music service The Discover (i.e. ‘home’) tab. The navigation is considerably more condensed (for the better) in the beta version of the app, and again, this also has Jon’s custom CSS on top of it.

Record labels: like Apple Music, they’re all clickable, which not only takes you to a label page, but also means they can be saved to your library, so they exist on the same hierarchical level as artists, releases or playlists — and can be followed. This is such a great way of discovering music.

Screenshot of the Qobuz music service The label page for Ghostly International — navigated to by clicking on the label name within a release page.

Magazine: With a strong focus on editorial content, Qobuz’s own digital magazine is built right into the experience. I didn’t see the benefit of that at first, but when magazine content related to artists started showing up on their profiles or on release pages, it all made sense.

Screenshot of the Qobuz music service The Magazine’s homepage. Screenshot of the Qobuz music service Note the magazine content showing up on PJ Harvey’s artist page.

Catalogue vs. library: If you’re on a release page that’s been saved to your library, and then click on the artist name, it’ll initially only show you releases from that artist in your library, and my first reaction was this is bonkers! Why would anyone want that? but then realised that you can then toggle to view the full catalogue, and this makes so much sense. This is exactly the approach Apple Music should’ve taken to overcome the confusion around local files, treating the release as the primary object, with metadata attached to it, rather than the other way around.

Screenshot of the Qobuz music service Left: Nils’ Frahm’s artist page with the “Catalogue” toggle on; right: the same with the “Library” toggle active. Screenshot of the Qobuz music service A close-up of the “Catalogue” / “Library” toggle on Queens of the Stone Age’s artist page.

Of course, Qobuz isn’t immune to metadata weirdness, and actually this is the only service I’ve noticed this on: all music ever made by Dominick Fernow is attributed to his main alias Prurient, rather than to his respective monickers, such as Vatican Shadow (Muslimgauze-esque techno) or Rainforest Spiritual Enslavement (dark ambient). It might seem like a nitpick, but I’m sure it’s not what the artist would want, given that he uses those distinct identities to release music across totally different genres.

Screenshot of the Qobuz music service This release should be credited to Rainforest Spiritual Enslavement, not Prurient (even though they were composed by the same musician).

Beyond this metadata niggle, there are some other points that work against Qobuz, too:

These irks are not small. But, in my opinion, Qobuz does a far better job at being a very capable music streaming service, with more solid apps across various platforms, than Deezer, Tidal, YouTube Music, Amazon Music, and yes, Apple Music. And it’s a family-run business, based in France, that clearly cares about music, paying artists 4× the industry average. In terms of simply trying to do something different, it’s hard to fault chez Qobuz.

There are certain things I miss deeply about Spotify. In fact, I’d go as far as to say that I think Spotify is actually the best music streaming service that currently exists. But if, like me, you feel morally uncomfortable paying for a Spotify subscription, then consider Qobuz as a worthy replacement that — perhaps surprisingly — is way ahead of the more established competition.

Read the whole story
emrox
5 hours ago
reply
Hamburg, Germany
Share this story
Delete

Play snail racing simulator online 🏁🐌

1 Share

Pick a snail and get comfortable. It's a long race and overtakes are rare, but the snails are taking it extremely seriously, so the least you can do is watch. Unless, of course, you happen to know how to encourage one.

Need to settle something?

The real thing

Snail racing is a genuine sport, and a wonderfully British one at that. Somewhere between a village fête, a sporting tradition and something invented because it was raining, people have been gathering to watch numbered garden snails cross damp cloth for decades.

The World Snail Racing Championships have been held in Congham, Norfolk, since the 1960s. Competitors are common garden snails, Cornu aspersum if you're being formal. Each gets a number painted on its shell and starts in the middle of a damp cloth. First to cross the circle 13 inches away wins. That's it. That's the sport.

The record belongs to Archie, who covered the course in exactly two minutes in 1995, a blistering 0.006mph. More than thirty years later, nobody has caught him. Heikki managed 3 minutes 2 seconds in 2008, Terri 2 minutes 49 the year after, and Sammy got it down to 2 minutes 38 in 2019. The dream remains alive.

Over in Cambridgeshire, the Grand Championship Snail Race at Snailwell has been running since 1992 and can attract up to 400 spectators, more than doubling the population of the village. The starting call, "Ready, steady, slow", came from the 1999 Guinness Gastropod Championship and remains the best thing anyone has ever said to a damp cloth.

The snails here race the same distance at roughly the same pace, with approximately the same understanding of what's happening. None of them has ever failed a drugs test. Mostly because nobody has tested them.

Know your snail

Snails are molluscs, distant cousins of the octopus, though you wouldn't guess it from watching them. They breathe through a hole in their side called a pneumostome, their blood is faintly blue, and they eat using thousands of microscopic teeth arranged on a ribbon called a radula.

Their eyes are on the ends of the long tentacles. The shorter pair are used for smelling and feeling their way around, which seems sensible when your top speed is measured in thousandths of a mile per hour. They don't have ears either, so they can't hear a thing, cheering included. Please cheer anyway. It's good for morale, probably yours. The shell is mostly calcium carbonate, the same stuff as chalk. A snail hatches with a tiny shell already attached and adds to it as it grows. Most land snails are both male and female. Given the choice, they'd do most of their moving at night or after rain, which goes some way to explaining the times.

There's more on Wikipedia, all of it apparently true.

There are snail racing achievements to be had, if you're the collecting type.

Read the whole story
emrox
5 hours ago
reply
Hamburg, Germany
Share this story
Delete

Introducing System One Models & Jev

1 Share

Diogo Almeida, founder, TypeSafe

Models have been superhuman at chat for years, so where is all the automation?

This has been my driving question for the last four years. At OpenAI, I helped build the methods that made language models useful at following instructions and talking with people. That work ended up as the research behind ChatGPT.  At the time, I thought maybe chat models would lead to AGI, but despite the hype it became obvious to me that there was something really big missing.

After two years in stealth, countless technical challenges, and research breakthroughs… I am beyond excited to announce that today, TypeSafe AI is releasing our first System One Model: a new class of frontier models built to make fast, structured decisions that software can use directly.

We built a new stack entirely focused on automation: with a new model architecture, parallel sampler for maximum efficiency, and training method we call Reinforcement Learning for Calibrated Decisions (RLCD).

Our first public model is Jev, available today in early access. Jev achieves similar levels of intelligence on System One tasks compared to existing LLMs, while being two orders of magnitude faster and more efficient. While Jev gives up string generation, it’s optimized for structured outputs and can’t hallucinate. 

Think of Jev as a frontier-intelligence function call: unstructured state in, typed probabilistic decisions out. 

Extraordinary claims require extraordinary evidence so see below for the receipts. 💅

Frontiers, Old and New

Existing LLMs

System One + Jev

Optimized with

Reinforcement Learning with Human Feedback (RLHF) / Reinforcement Learning with Verifiable Rewards (RLVR)

Reinforcement Learning for Calibrated Decisions (RLCD)

Optimizes for

Human preference: writeups and chat responses that human raters prefer.

Verifiable rewards: outputs that can be programmatically verified.

Calibrated decisions: answers with epistemically honest probabilities on System One tasks.

Inputs

Unstructured data (e.g. text) with an emphasis on sequential messages.

Unstructured data (e.g. text) with an emphasis on structured program state.

Outputs

Strings / generated text. Strings are flexible and can be anything: chat responses, code, hallucinations, refusals, or even type-safe structured values. To be used by software, responses need to be parsed + validated. There is also always some risk that the AI goes off the rails.

Type-safe structured values. Possible outputs and structure are defined in advance. The model never makes type errors. All answers are accompanied with calibrated probabilities and confidence scores.

Sampling

Sequential. Generates one token at a time, each conditioned on the last.

Parallel. Generates all outputs in a single query. Incredibly efficient and hardware-aware.

Cost

Input tokens: from $0.20 to $10 / MTok.

Output tokens: ~5x more expensive than input tokens.

Input tokens: $0.042 / MTok ($42 per billion tokens).

Output tokens: FREE (too cheap to meter).

Speed

End-to-end response time is 3 to 329 seconds for frontier models.  Fast enough for interfacing with humans, but a big bottleneck when integrated in code.

End-to-end response time is 70ms-500ms for TypeSafe. This can range from 40x-200x faster for the same levels of frontier intelligence for System One shaped queries.

Confidence

Even if prompted for a confidence estimate, models tend to be overconfident and inconsistent. If a model can do a task 95% of the time but doesn’t say when it’s in the 5%, it can’t automate that task.

Always communicates confidence and uncertainty with every output. Calibrated: higher confidence means higher accuracy. More consistent: returns similar answers for similar inputs.

Use cases

Human-in-the-loop tasks (chatbots, copilots, coding agents). General and powerful, but requires human oversight because their freedom also means they might go off the rails.


Verifiable problems (math proofs, kernel optimization). When correctness can be checked cheaply and automatically, LLMs can generate, test, and iterate until they find something that works.

Demos. The flexibility of strings allows it to be incredible for quickly making prototypes that only work sometimes.

AI-Powered Workflows / smart if-statements. Structured outputs slot into ordinary software as fuzzy decision rules: classify, route, score, extract, or branch where hand-written logic is too brittle. The surrounding code constrains their freedom, making them easier to compose into reliable systems.

Map-reducing over big data. Turn petabytes of data into features and insights.

Real-time applications. 100ms speeds means you can use AI in your applications where UX is critical.

Verify everything.

Score, judge, verify, guardrail, and detect jailbreaks of LLM prompts, reasoning traces, and/or outputs.

Evidence / Technical Results

We love skeptics, and are skeptics ourselves.

There are some claims you can easily verify:

  • Speed per call: We truly are that fast, though our published evals are generally run from our laptops on the West Coast (this is where our service is currently based).

  • Cost per call: We make our pricing transparent. We can’t prove it isn’t subsidized; we’ll need the long-term to prove the sustainability of our pricing (which we expect to go down, not up).

  • No type errors: This would be an easy thing to falsify with just a single counter-example, but it is mathematically impossible.

For our bolder claims, we want to provide as much nuance as we can.

Side-by-side demonstration

Our side-by-side demo shows a key difference between our models and LLMs: Jev outputs all probabilities in parallel instead of autoregressively generating by token. Strings are extremely powerful and general, but costly. “Giving up” strings actually gives us a lot of superpowers!

Nuance
  • For people with early access to TypeSafe, here is the actual query.

    • The query is highly simplified and questions were chosen to have descriptive, human-readable keys so that the output on the screen is understandable.

    • The state is also a short, dense, and detailed paragraph, to emphasize the difference in sampling methodology. The relatively shorter input paints our model in an advantageous light.

  • For the keen eyed, for the recorded run, the only disagreement with GPT-5.6 Terra is on “Churn likelihood level”. The actual answer seems genuinely ambiguous to us.

  • We used GPT-5.6 Terra with default reasoning for this example, because we’ve found it to be the most comparable at intelligence to Jev on average.

  • Fun fact: a similar demo was what convinced us to go all-in in the direction of System One Models!

Workflow evals

We made a new type of evaluation to measure how well AI works within code. We don’t optimize for a ground truth classification or allow the harness and model to change (potentially allowing for overfitting via harness engineering). Instead, we assume there is a correct compute graph (a “workflow” represented in code) and use the predictions of the largest, smartest, and most expensive external models as reference probabilities.

Rephrased: every model gets the same workflow. We test how they compare to the average of the smartest models (in this case, Astra and Fable).

Jev is off the charts – owning the Pareto frontier for almost 2 orders of magnitude. We also compare to models with a generated prompt doing all the logic in their chain-of-thought, but this tends to do significantly worse than using the workflow itself.

Note that the calls here are significantly more complex than the side-by-side demonstration above. That’s because they’re more representative of the types of production workloads needed for true business automation. Below is the simplest of the 4 workflows we’re publishing:

The most reliable real-world workflows tend to have many independent, decomposed questions, with fine-grained behavior that’s dependent on probabilities instead of discrete decisions. The end result is discrete branching, but how we get to a final answer involves a lot of domain-specific engineering that needs to be done highly consistently.

See our workflow evals site for all the details: examples, disagreements, full queries, and each workflow.

Nuance
  • This is where the claims of 193.6x faster, 444.6x cheaper on our home page comes from, and we expect that these are on the higher end of real world gains.

  • These content of these workflows were not deliberately chosen nor constructed to make our model look good, and are not in our training distribution. However, they were made by individuals on our model capabilities team, so some bias could exist.

  • We use the average of GPT-6 Astra and Fable 5.1 as the reference answer, which biases answers towards OpenAI and Anthropic’s models. We likely underestimate the relative performance of our model and DeepSeek’s models.

  • The LLMs use our System One LLM wrapper, which constrains LLMs to output structured decisions compatible with our API. We have found this to be the most accurate way to get decisions from LLMs, but this tends to be slower and more expensive than giving decisions without probabilities.

Hallucination and Type-safety

Hallucination and type-safety are intrinsically related, and we think the latter is table stakes for automation. Having a hallucinated tool call is inconvenient in an agent, but is an absolute deal-breaker if it’s part of a system with latency guarantees or it’s buried several layers deep in a dependency chain. Existing models, no matter how smart, still hallucinate and have type errors.

Nuance
  • The numbers for LLMs are from OpenRouter i.e., there almost certainly is bias here: more complex queries might be routed to better models.

  • Our number is not empirical. Schema matching is guaranteed, thus we can confidently add 0% into the plots.

Fun Demos

Perhaps the most exciting part of our work is enabling new use cases. We have a lot more to show you, but here are a couple of the team’s favorites:

Doom

We love how this doomo doomonstrates real-time intelligence and what can be doone with code + AI. The engineer behind it was worried about making 10 queries a second (which ends up costing ~$7/hour), but the rest of us agreed that was lower than expected! This is so fun we intend to not only release an in-depth walkthrough, but also host some events to hack on this.

Nuance
  • The demo is on structured state as a data structure with text, not on images (yet…)

  • A non-AI doom bot could play better, but we wanted a bot that was reactive to different representations of game state, and most importantly… following instructions was cool as heck!

Wikiracing

The objective of the game is to start on one Wikipedia page and reach a specific other Wikipedia page using only links you come across while traversing. Each step can mean choosing between hundreds to thousands of links! It’s a great playground for demonstrating not just intelligence-per-second, but also the compounding benefits of not hallucinating with high-cardinality choices.

Nuance
  • As far as we know, it was completely random that both the 2nd and 3rd challenges started with “Rubber Duck.” The author only noticed when the team pointed it out.

  • Our speedups here tend to be a lot less than in previous demos. That’s because this is against the non-reasoning modes of the models (except Astra which was set to the lowest reasoning setting). This is also why Jev tended to finish in fewer steps (a sign of greater intelligence). This was to make the demo more bearable to watch. The LLMs look much worse at this task than with reasoning enabled.

  • Jev supports a cardinality up to 255. For the higher cardinality choices, we do a 2 stage-system of scoring independently then making an explicit choice, hence the occassional slowdown.

What’s next

We’re still in Jev’s early days. We have a lot more in the pipeline and are so excited to keep on shipping 🔥.

Today, we are opening early access and bringing developers off the waitlist as quickly as we can. We want to hear which decisions you need to automate, where Jev works, and where it falls short. Tell us what sci-fi you want to build!!

We started TypeSafe because we believe that AI needs an interface software could depend on. We can't wait to see new use cases continuously diffuse through the community and economy.

We Give A FAQ

Where do the names “System One Models” and “Jev” come from?

We were inspired by Daniel Kahneman, Thinking, Fast and Slow. The model class name draws on the distinction between fast, intuitive System 1 thinking and slow, deliberate System 2 reasoning.

“System 1 thinking” has also implied error-prone. For reasons we will get into in the future, we believe System One Models can be made more reliable than its alternatives.

We named Jev after William Stanley Jevons. We expect machine intelligence to follow a similar path to coal, after steam-engine efficiency led to an increase in demand. Every order of magnitude drop in the cost of intelligence unlocks orders of magnitude more use cases.

Why was a new training algorithm needed?

What use cases is Jev good for?

Is Jev just a smaller LLM?

How does Jev perform against public benchmarks?

Where does our training data come from?

These are results are kinda crazy - how is it possible?

Read the whole story
emrox
7 hours ago
reply
Hamburg, Germany
Share this story
Delete

Making Your Data Ready for Agentic AI

2 Shares

Lots of organizations are excited about what AI can do to streamline their processes, save money, and juice margins. But AI's capabilities are founded on the data that AI accesses, and for many organizations that foundation is little more than sand. Pramod Sadalage and Prem Chandrasekaran write about how to build a reliable foundation of data that can be accurate and trusted.

more…

Read the whole story
emrox
8 hours ago
reply
Hamburg, Germany
alvinashcraft
24 days ago
reply
Pennsylvania, USA
Share this story
Delete

HEIF Heist

1 Share

Hacktron AI

One image parser to pwn them all

01 Overview02 Research origin03 FAQ

What is HEIF Heist?

A bug that could have allowed us to

HEIF Heist is Hacktron's name for a class of remote attack paths targeting services that decode attacker-controlled HEIF, HEIC, or AVIF images. By exploiting underlying native libraries, these vulnerabilities allow an attacker to bypass application-level defenses and trigger memory corruption, data exposure, or remote code execution (RCE).

The vulnerable attack surface lives below the application layer inside native C/C++ decoders such as libheif and libde265. These parsers typically enter production environments indirectly bundled via higher-level wrappers like ImageMagick, libvips, or Sharp, standard distro packages, and prebuilt container base images.

By probing upload endpoints with crafted .avif or .heic files, an attacker can fingerprint the remote libheif version family in use. Once identified, they can fire an exact version-matched n-day or 0-day payload to trigger memory corruption, data exfiltration, or remote code execution.

Research origin

A precarious tower of stacked dependencies, each block resting on the one belowEverything up top is resting on something underneath.

HEIF Heist began as part of the Hacktron research team's broader security research into frontier labs. After discovering and reporting a libheif RCE in Discourse, we asked a larger question: how many other applications depend on the same image-processing stack?

Past vulnerabilities such as ImageTragick, ForcedEntry, and the libwebp flaw have demonstrated the reach of an image processor or parser vulnerability. An image parser might generate an operating-system thumbnail or process a web upload, giving it an enormous blast radius.

That initial finding grew into a multi-month investigation tracing libheif across communication platforms, cloud services, enterprise products, and popular web frameworks.

FAQ

Hacktron

Work with the team behind this research.

Hacktron brings together top CTF researchers, experienced red teamers, and offensive security researchers. We use AI to accelerate security research, finding and eliminating vulnerabilities in widely trusted software before malicious actors do. We're continuing our research across frontier labs and other internet-critical systems. If you're responsible for securing one of them, we'd like to work with you.

Book a callExplore Hacktron

Read the whole story
emrox
9 hours ago
reply
Hamburg, Germany
Share this story
Delete

Audit your Agent files

1 Share
A practical guide to auditing what your coding agent still needs.
Read the whole story
emrox
12 hours ago
reply
Hamburg, Germany
Share this story
Delete
Next Page of Stories