Episode 17

Small Models = Content Moderation at Scale

with Dave Willner of Zentropi

Show Notes

My guest today is Dave Willner, cofounder of Zentropi and a fixture in the Trust and Safety world. In 2008, he was a young Facebook employee who sat down to replace a one-page list of banned content — it said, roughly, no Hitler and no naked people — with the first real set of community standards the internet had ever seen. The rules he wrote are still the foundation of how Meta governs content shared by billions of people. He went on to build community policy at Airbnb, and then joined OpenAI in early 2022, before ChatGPT existed, to figure out how to stop an image generator from producing bad things.

Now he's co-founded a company called Zentropi, with Samidh Chakrabarti, which is built on the idea that moderation at scale was never actually impossible — we just didn't have the right tools.

Some themes:

  1. Why human moderation breaks down — Human moderation suffers from high turnover, trauma, and inconsistent judgment — problems small, tailored language models solve by strictly enforcing written rules.
  2. Small models, big advantages — Specialized small models can run circles around massive frontier models on speed and cost, making real-time classification financially viable.
  3. Policies and labels are the same coin — Written policies and labeled datasets are two sides of the same coin; when an AI disagrees with a human label, it usually exposes a vague policy rather than a model error. Running moderation models locally on your own servers cuts per-check costs, dodges European privacy paperwork, and keeps user data in-house.
  4. Adversaries move at AI speed — Bad actors are using AI to scale exploits at high speeds, forcing safety teams to adopt AI defenses immediately just to keep up.
  5. No one-size-fits-all rulebook — A single global moderation rule doesn't work — dating apps, teen platforms, and social feeds require custom enforcement thresholds based on their specific risks.
  6. Lazy evaluation loops — Using the same LLM to draft a policy, generate test cases, and grade itself creates lazy evaluation loops that hide major safety blind spots.
  7. Power to volunteer moderators — Giving open-weights AI tools to volunteer moderators on Reddit or Discord lets small teams govern massive communities without burning out.

Chapter Timestamps

  • 00:00Intro
  • 3:21Dave Willner's trust-and-safety background and the origins of Zentropy
  • 7:04Zentropi's specialized moderation models and policy-optimization tools
  • 11:15Why human moderation is inconsistent and difficult to measure
  • 15:45The limits of human judgment and the case for automated moderation
  • 21:09Imperfect language, model precision, and faster adaptation to abuse
  • 24:43Shared moderation infrastructure versus platform-specific policies
  • 28:46Why small specialized models can outperform larger general models operationally
  • 32:26Local deployment, privacy, and trust-and-safety resource constraints
  • 34:16Customer economics and real-time video moderation
  • 38:09The danger of “slop” policies and weak AI evaluations
  • 44:10AI moderation as a tool for decentralized community governance
  • 47:06The adversarial threat landscape and urgency of AI-enabled defense

Transcript

There may be transcription errors: we apologize for those in advance.

Rob: My guest today is Dave Willner. In 2009, he helped Facebook write its first set of community standards. The rules he wrote are still a foundation of how Meta governs billions of pieces of content posted daily. He went on to build community policy at Airbnb and then joined OpenAI in early 2022. Two years ago, Dave co-founded a company called Zentropi, which is using small language models to help companies do content moderation at scale. I hope you'll spend some time with us diving into this. I really enjoyed the conversation with Dave and as always, please send me your feedback. Hey Dave, how are you?

Dave: I'm good. How are you, Rob?

Rob: Doing well. I'm doing well. So I saw you at TrustCon a few months ago. When was that?

Dave: Gosh, two months ago? July 20th to 22nd, so month and a half.

Rob: And that was my first TrustCon. It was a great event. I really enjoyed it. This summer, it's been like some family thing. It's been a conflict for the last few years, and then this year it was great. So what a great event.

Dave: Yeah, no, it was awesome. It's amazing how big it's gotten, and that was the fifth one. And the first year was like 250, 300 people at the Sheraton in Palo Alto. Very different vibe. It was kind of like a weird outdoor pool party because it's like this very mid-century suburban hotel. But they moved up to the Regency the next year, and it grew very rapidly. And it's now sort of, I think, reached the carrying capacity of the hotel, but also kind of reached like its stable equilibrium in terms of size, which is good. There's not really a bigger... There's only one bigger hotel in San Francisco they could move to, and it's much bigger and a weird space, so.

Rob: Yeah, that hotel. So I stayed in that hotel once. It was my first job out of college. Yeah. It was a very distinct memory. Stayed in that hotel the night before, we're supposed to essentially... I was working in financial services at an investment bank, and they sent us out to go and essentially sign up a client, take them an engagement letter, sign them, sell the company or whatever it was, and I was 23 or 24 years old. And so we stayed in that hotel the next day, go to the meeting, go to the client meeting, and the client's like, okay, I want to hear about your capabilities. And we're like, oh no, this is not signing an engagement letter. This is a pitch. We had no materials, and you know, like, investment bankers are like famous, we're like pulling all-nighters to create pitch books and all kinds of nonsense. We've done none of that. And so the person that I was working with went on to be a well-known investment banker. I did my two years and was out, but I saw this, we just stumbled through this thing. I knew nothing, wasn't expected to know anything, but this was like one of those crazy professional experiences where you're like, we have no idea what we're doing. This is going to be extremely painful, and we're going to move through it. And I think the segue here is, you know, like content moderation. We have no idea what we're doing. We're making this up as we go along, and we're like in this meeting, and it's probably going to go, you know, I don't know where it's going to go, but it ended up going nowhere, but you know.

Dave: Yeah, it does sound like good prep as a professional thing-maker-upper.

Rob: So by way of introduction, Dave Willner, what are you working on? What were you working on before? Like, what's your professional story?

Dave: You want the whole biography?

Rob: Give me like the capsule version.

Dave: So, okay, I've worked in trust and safety for the better part of the last 20 years. I started in the field in 2008 as the 12th content moderator at Facebook doing the actual moderation, looking at all the blown-off heads and naked people. And at the time, there wasn't really much of a plan. We didn't have any policies, we didn't really have much in the way of tooling. And so I ended up sort of volunteering for a project to write a systematic set of policies. But as we have all learned in the 20 years since, policies are always and forever imperfect and never finished. And so that project turned into a part-time job, which turned into a full-time job, which turned into a team of people building the organic content policy. And as part of that, was also pretty involved in a lot of the sort of work on the tooling to enforce the policies, which I think has informed a lot of my subsequent practice, because I don't really think these questions of operations, tooling, and policy are separate questions. They're just kind of different parts of the same elephant and viewing them in isolation is in my view usually a mistake. So did that for like six years, left, went to a small startup that failed, then ended up at Airbnb running their community, building up their community policy practice because they had gotten themselves in a similar situation where they had a bunch of enforcement, but not a ton of codified rules. Was there for about six years, again, worked on tooling, also ran the quality and training teams for a while for Trust, left, was at a small startup that failed. And then went to OpenAI running their Trust and Safety team for sort of the applied startup be part of the company when OpenAI was still mostly a research lab with a pet startup instead of the thing it has become. And was there running that for the launches of DALL·E, ChatGPT and GPT-4 for between the consulting part and the advising part and the employee part, like a couple of years, and then ended up at Stanford studying how we might use LLMs to solve some of the intractable challenges that I had been failing to solve for the last 15 years before that. And met Samidh Chakrabarti, who I knew of but had not ever met, who had come to a similar conclusion through his own sort of path through this strange world of Trust and Safety. And we started collaborating on actually pulling together practical ways to make a small LLM good enough at following custom policies that you gave it to actually be able to use it in production because you can do stuff like that at the time by bullying the default models, but you couldn't afford to. They were too slow and too big. So the trick was, can we teach a small model to do this really? Well, the answer turned out to be yes. And we then decided, oh, you know, we want to work on this full time. So spun it out as a startup, which is what I'm doing now at Zentropi.

Rob: That's awesome. So how long has Zentropi been going for? And yeah, tell us a little bit more about what you've learned so far.

Dave: Two years this Halloween, it'll be two years this Halloween, our incorporation date is Halloween. We didn't tell everybody we existed until December of that year, but technically the incorporation date is Halloween. I do a big Halloween, elaborate Halloween set up. And so Samidh's plan is that eventually that becomes the corporate party long term plan. It's going really well. We work on a couple of different things. So we still work on special purpose, small models for sort of bring your own policy classification. We're at this point on our second generation and working on our third. The model is expanded from, you know, a relatively small context window, along with sort of the advances in small models that we're training on top of it now, you know, accommodates like a couple hundred thousand tokens of context and is actually faster and can do images and can do video and all of this stuff. And then we also build agentic systems that take advantage of a model that can do that work to build sort of unique systems to help you use it. So to be more concrete, if you've got a policy following small model, that's great. It lets you update your policies really quickly in response to changes in your user behavior, right, but you're still you're still bad at writing policy like writing policy is still hard and honestly historically was a bit more astrology than science because rolling it out was so laborious. You paid someone like me to basically guess what your moderators would do in response to the things that they wrote. Like that was my job. Right. It's like it's like moderator-reading astrologer and you're writing a set of instructions to try to get people to do what you wanted them to do based on how you assume they're probably going to read it based on your understanding of how they misread things. And we sort of figured, hey, if you've got this model that responds instantly to changes you make in the policy, you can actually just start to empirically test how to say things to do to get the model to do what you wanted to do. And so we've built this set of agentic loops that can both take your data set and your policy and give you back a better articulated version of the policy and can return to you places where we think your canonical labels that you say you stand behind are probably wrong, at least if you are holding yourself to the policy text you gave us as the standard for what is correct.

Rob: Just kind of like driven like drifting away from your intention. Is that the way to think about it?

Dave: Another maybe a core premise to all of this is that superhuman in the AI space sounds very impressive because it has the word super in it. But if the humans are decidedly mediocre, it's not that high of a bar, right? And it turns out that people are pretty mediocre at large scale content labeling for because it's boring, because it's traumatizing, because it's not a fun job, because it's not well paid, because consistency over time is remarkably hard, because it's actually a very cognitively difficult task. There's a bunch of reasons, but just about anybody's golden data is like internally inconsistent. They're not really fully consistent with how they're labeling that data. And just about anybody's articulated written content policy doesn't like really fully capture exactly what they meant, because both of these processes have been very sort of artisanal and squishy traditionally. And so if you're going to try to have a robot optimize those things, you actually have to optimize the golden data set and the policy together, because they're really a single object. There's no such thing as a golden data set aside from the policy that defines it in our way in speaking, because this whole thing is a social construction project, where you're asserting a set of rules that decide what things are for the purposes of this exercise. And you can view the labeled data and the policy as sort of two imperfect artifacts of your intentions and your goals. And what we do with our machinery and are trying to figure out how to build is stuff that looks at those figures out where they agree, which we take as a signal of your strong intentions, and figure out where they don't agree, which are gaps to close in either how you articulated what you meant, or confronting you with things where it seems like you're maybe not sure what you wanted.

Rob: Do you find that in your work so far, do you find that the customers or potential customers understand this distinction? Or where do you feel like people are? Because I have mostly some thoughts on what the perception of humans versus nonhumans in various kinds of moderation might be. But do people kind of get when you're talking about this, do people kind of get that distinction, do you think?

Dave: Yeah, I'm usually less philosophically abstract when pitching people on the system itself.

Rob: No, this is good, though.

Dave: Yeah. But yes, they do. And it's partly because the system itself helps reveal this, right? Like, if you bring us a set of golden data, and we label it, and we come back with using the model and your policy, and we come back with some F1, it's not actually that abstract to say, like, okay, well, here's the cases where our model says something different than what you said in your golden data, but like, who's actually wrong here? And that becomes a comparative review process where you're like, okay, well, you may not have wanted the outcome that the model is providing, but that's not the question. The question is, what are the instructions you gave dictate? And if the instructions you gave dictate something different than what you marked, but do dictate what the interpreting model marked, well, either you marked the wrong thing, or you need to clarify what you wanted in what you wrote down, right? It's one or the other. And that's not to say the machine is not perfect, it sometimes is also just wrong, right? There is a set of errors. But I would say that in practice for most folks, outside of the very largest companies that are really, really sophisticated about this, their data set labeling practices are not high quality enough that the vast majority of errors are not actually just misalignments where maybe it's what they wanted. Maybe it's not what they wanted, but it's not, it doesn't follow from what their policy said. And so just confronting folks with those examples is a pretty decent way to get people to understand this distinction. The other thing I'd say is, there are a lot of companies out there that can't have the sophistication of a Meta or a TikTok or a Google in terms of their human operations, and they know it. And particularly if they're AI-first companies that are already inclined to believe that you can build interesting and new things with AI technology and don't have a lot of existing human operation systems built up, that's a much easier conversation. Or if for resource reasons, they can't be doing that's a much easier conversation. So it really depends. It depends on who the customer is, where they are in the process, what their own systems already look like. It's all over.

Rob: I mean, is it fair to say that a lot of policies for content have basically just grown up or somewhat organically out of necessity and out of experience. And that's been a very imperfect process. Now you finally have the means, the mechanisms to actually add more rigor to that, even if you're not a huge corporation that can throw resources at the problem.

Dave: That's kind of where we're at in this cycle. Yeah, that's exactly right. You can be more, I'm not going to say scientific because that's too much of a claim, but you can definitely be more empirical in the sense of, because a labeling model responds instantly to the changes you made in the text, you can just test how to explain your feelings to get the result that you want. And there are better or you can produce explanations here being policies. You can produce guidance that does produce the results you want, but are poorly formed. So there is still skill and art here, because you can just stick all the examples in the policy. That's not a good policy though. It's a form of overfitting in a policy text. There are still better or worse ways of doing that, and a bunch of our systems help you do that fitting in a way that we view as resulting in good policy instead of resulting in sort of cheating or sloppy or overfitted policy. But it is the case now that we can just like test a lot of approaches and just pick the one that actually works instead of historically where we were, where you really couldn't do empirical approaches to policy text drafting.

Rob: So one of the things that I always found fascinating in conversations with journalists and others and regulators was this idea that if you're not having humans look at stuff, you're not trying hard enough. Because they didn't have empirical ways to assess how good you are at making various actors follow your policies or enforcing the policies you've written or having the right policies or any combination of those things. And so turns out, hey, we want humans to do this stuff, but there isn't necessarily a sense of whether or not humans are good at content moderation. So is it fair to say you think humans are not good at content moderation?

Dave: I think humans are not good at high degrees of precision at large scale, kind of in anything. Like, I don't know, can you draw a circle freehand? I can't. Like, my mom can she has an MFA, right? Like she went to she has a master's degree in being able to draw a circle. Well, like, and then that's not to run it down. It's me saying that's actually incredibly hard. Like, we have many unique and wonderful gifts as people and precision is not typically one of them outside of very specialized work at a relatively slow pace. And so if you need to make a lot of something very precisely, I think actually in any domain, the one answer we as a civilization have ever come up with is figure out a way to get a bunch of robots to do it so that you can tune how precise they're being. Like, that's the answer to this question. The claim I'm making here is that LLMs mean that now that kind of thinking, that kind of industrialization can start to happen to what was already a factory process, right? It was just a factory process with no machinery at any stage in the production line, which was this sort of content labeling system. I do think that sometimes strikes people as a radical view. It was honestly influenced by some of my early time at Facebook. Before I got there, Andrew Cuomo, weirdly, was the Attorney General of New York. And the 50 states had sued the company about turnaround times on nudity reports and reports of abuse. And the company settled. And as part of the settlement, they agreed that a human person would make all of the decisions on nudity reports and all of the decisions on emails to the open abuse alias that we got for the next several years and that we would be audited by the security company, Kroll, to ensure our compliance. And so when I say my job was looking at the blown-off heads and naked people, it was literally like we have to look at these naked people in 24 hours and we're not allowed to use any machines to help you or we're going to go to New York prison. And even in like 2010, that was bad. It was like working on the inside of it. It was obviously bad, right? There were things we could have been doing to triage away spam emails to the abuse at email box, which we weren't allowed to do because the spam filter was a robot. And that's just dumb. That's just like clearly dumb. And so there's like a deep wound for me that I think influences this perspective. But also like it comes out of trying to get high precision from quality, from running quality teams, from running training teams. We're all very different. These decisions are very complicated. And to your point about the sort of unmeasurability of problems, a lot of the worrying that I hear in the AI space around potential bias, which is a real thing we need to care about, or potential mistakes, which is a real thing we need to care about, or changes in language and terminology moving quickly, which is a real thing we need to care about. Like, all of the objections I can log are also true of making 1000 people do it. Except in the context of making 1000 people do it, you're also causing a bunch of suffering to that 1000 people. And the 1000 people are constantly changing because it's a terrible job and they turn over every six to nine months. So you don't even have a static body of 1000 people whose bias you can begin to attempt to measure. Right. And so to me, taking a chaotic system and moving it into a measurable system at least means we know what is wrong and can maybe start to attack the ways in which the system fails in a systematic way, instead of the prior alternative, which to me feels like very romantic, which I mean, I think is what you were hinting at, it feels like sort of very romantic and comforting to believe in this notion of like an ineffable shared human spirit, which knows the right when it sees the right. But like, that's not a thing, guys, I don't think that's a thing. Like, I think we all have intuitions about what's bad and what's good. But how we implement those in specific decisions is very nonuniform. And everybody wants uniformity out of these decisions so that the company can, so that whoever is making the judgments can be said to have a standard that they're treating people fairly. So I guess

Rob: I guess two questions on the uniformity. So, you know, do you think that so if humans aren't necessarily going to do content moderation, well, can machines do content moderation? Well, I think is the first question. And I'm going to hold the second question, because it'll follow naturally from your answer to the first question.

Dave: I think there's a limit to how well you can do content moderation, which is where my industrial analogy falls apart, because words are not steel. The carrying capacity of precision of language is just less than other things. Like, this is why, you know, formal logic notation exists. This is why math exists, right? Newton invents calculus, and he's writing these long ass paragraphs. And Leibniz says like, this is dumb as hell, let's invent a different notation to do this. And then Newton bullies him about it and all of that stuff, right? Like, there is some limit to the precision of moderation, from the fact that this is all a word game, right? And that ultimately, all of these meanings are things we make up. And this whole thing is a social construct. And even the language we're using to talk about the language we're creating violations in is a social construct. So like, yes, we're never going to get to perfect moderation. That's not a thing that can exist. To me, that's just a given. I don't that's not a solvable problem. That's a fundamental, just like challenge or not even challenge. It's a fundamental bound on the entire exercise, no matter how you do it. The question to me is what percentage of the stuff that goes wrong in moderation, you think comes down to imperfect execution within those limits versus some sort of imperfect will or a lack of trying hard enough. And I think the public narrative around moderation has typically been moderation is going badly because the companies don't give a shit and they aren't trying. And look, I'm not going to say that some of the companies have not made decisions I disagree with in terms of the actual standards. But like, as a thought experiment, if you were Mark Zuckerberg, and you ordered your human moderation system 10 years ago, to perform exactly your preferences, like we couldn't have done it would be it would actually be more comforting in my view, if the problem were a problem purely of intention and effort. Because then you could like, sue everyone and get different CEOs and pass some laws and stuff. And then it would be better. And the version of this story before the rise of LLMs I used to tell ended much more darkly. It was me saying like, this guys, this is the best we got for you. We spent billions of dollars and a lot of really smart people bashing their heads against the wall. And like, yeah, sometimes there are things we've not been allowed to do. But frankly, those mostly aren't in the moderation realm. They're more in the realm of changes to things like feeds and algorithms where like, I totally with you. But on the moderation side, everybody's trying as hard as they can. And like, this is what we got. And it's never going to be any better. And I do think LLMs mean that whatever percentage of the problem is sort of a precision execution scaling problem, we can shrink that gap not to zero, but a lot. And in addition to the better precision, we also can adapt more quickly, which is sort of the other problem because it's an adversarial space. And the human power version of the system is incredibly slow to change. So it's both like not very good at precision and very bad at responding to adversarial events at scale. And the models are obviously going to be better at both of those as we figure out how to use them.

Rob: Yeah, so this leads me to the second part of the question. you kind of started getting me to the point of the second part of the question, which is that, does this just mean that there's real economies of scale or whatever in this area? And that you should have one model, you know, you should have like a single content moderation model that can apply for 99% of the use cases, as opposed to each company, each entity having, you know, their own take on it, right? Like maybe this is a naive or stupid question, but I really do wonder about this, which is like, should everyone be coming up with their, you know, organically grown or informed by other, whatever policies, and then essentially learn all the same lessons like many times, or could we just jump to a later state by having shared infrastructure between entities that have to do moderation?

Dave: Yeah, I think the answer to this is both yes and no, in that I think there is a set of problems for which the thesis you're posing is basically true. There's a set of things that are like crimes, or otherwise very highly socially condemned, that basically there's not a ton of point in reinventing the wheel. And I also think even in specific domains where there's more variability, like nudity is a great example, right? Like the nudity standard you're going to have for a dating app is going to be different than the nudity standard you have for for Facebook, or like, I don't know, Facebook for Instagram for teens, or whatever, right? Like, those are going to be different standards. And that's fine. And I guess that sort of hints into my second thing, which is, there are shared ways of defining things and talking about stuff from a policy point of view that, you know, we publish a bunch of the policy stuff that we come up with ROOST, which is the Robust Open Online Safety Tools project, which is a sort of open source trust and safety tooling project, hosts a bunch of policy stuff, some from us, some from other people, we think that's great, we think folks shouldn't be starting at ground zero. But it remains the case that you are essentially fitting a classifier to the distribution of problems you have. And because these detection systems remain imperfect, the right trade off is a function of both the line you want to draw in your sort of dating app versus Instagram for teens. And the distribution of content you have in your platform based on the users you have and the ways in which a given product allows abuse to occur. So like, I don't think you end up with one standard for everything because the right ways to be wrong, the right kind of mistakes to make in your system change based on whether or not those mistakes really matter in the system you actually have based on how data flows through it, right? Because there's like a little bit of a shell game we all play with classifiers where you everybody's arranging their failures, and you just try to put the failures on this category of things that doesn't actually happen very much in reality in your platform. And then you can like artificially drive down your failures. And that game exists in this world no matter what. And so I think you have to continue doing it. And then the other thing is like people come up with new and spectacular ways of being awful constantly. And so there's always going to be learning. We don't offer this as a service yet, because we're just training the best version of an interpreting model, we know how to train. But I do wonder if there's a future where there's multiple versions of these interpreting models that have sort of slightly different biases and slightly different habits of mind. And if you end up with something like multiple reviewers, because you're passing the same policy through multiple different interpreting models, and getting different reads, and this people talk about this as a as like a model chorus or a model panel of judges. And I do think you're starting to see some of that emerge. Right now, it feels like it's mostly a hedge against the fact that a lot of the models people use to do this are not specialized adjudication models, they're just mid size foundation models. But I do find myself wondering if there's a world for that even in a future of multiple specialized bring-your- own-policy classifier model.

Rob: And that makes a lot of sense. I think the point you made earlier about, you know, having the multi purpose model that is slower, it can do a lot more, but like maybe is overkill for your use case, I think it just makes sense that people would use a lot more specialized stuff. Well, specialization gets you speed. And we all

Dave: know this, right? But like, speed is very powerful, right? Being able to do eight times as many classifications, 97% as well, is like pretty great, right? Almost everyone would take that trade. And so the part of our magic trick here is also that CoPE is very, very, very good, like frontier competitive while being very, very cheap and fast to run. And you can simply build different systems with a system that is good and feel like quite good and very fast than you can with a system that is marginally better, but 10 to 100 times slower.

Rob: Yeah. So one of the things I hear from companies, there's some companies, they're building their own systems, they're using, in some cases, they're using frontier models, like the bigger models, that I hear the complaints about, oh, it's slow or, but you know, it's good, but it's slow, or I can't put it in the loop, I have to run these analyses offline or whatever. And, but I guess my question is, like, do those more generalized approaches just keep getting better or getting optimized or whatever? Or is that just fundamentally not going to work for most of these cases? Like, is it like, how do you have, how do you address that?

Dave: It's a good question. I wouldn't venture to predict the future there. I am not an expert. And I think future predictions in this domain are basically all foolish at the moment. It does seem to me, the big models keep getting bigger. And yes, we have more hardware and tricks to run them faster. But like cost is a real thing. And compute constraint is a real thing, even if we're not in the crunch that we're currently in. And so like all else equal, the cheaper, faster one is always going to be cheaper and faster. Right. And so my personal intuition is that there is a place for specialized, very quick models that do certain specific kinds of things extremely well, because you may need to do a very large number of those things. And I actually don't think that's true just in the content classification case, right? You already see small, locally runnable models in other areas of LLM development. There's a really interesting model called Cactus Needle. That is a tool calling only model. It's from a YC-backed startup called Cactus. It's like 26 million parameters, but you can speak to it in natural language and it'll do two tool calls, like against a list of tools you load it. And the thing is tiny. And so therefore incredibly fast. And you can put something like that on different devices and use it at different scales and afford to use it a different number of times. And like, I don't know that just seems like math to me, like no matter how many GPUs we have resource allocation exists. And so I think that space abides. The more interesting question to me isn't actually the really big models. It's whether the general small models get so good that specialization in the small domain becomes less relevant there. That hasn't happened yet,

Rob: but I don't know. I think trust is the other factor, right? Like, I think a bunch of people are distrusting of the bigger companies, or at least I've, I've heard this repeated a few times with that. Like I want to run this stuff myself, even if I'm giving up certain attributes or aspects of it. I'm assuming that's something that you see from the descriptions. Yeah,

Dave: it's been a huge draw. That's definitely been a huge draw for us. We let folks run an open-source, text-only model. Our multimodal model, which handles images and video, is closed, in the sense of it's not online for you to download. We let customers deploy it on their own infrastructure. Our API is logless, so we do not see your data there either. So like, we do not see your data there. But you are going to another server and there is the round trip costs and all the rest of that. The vast majority of customers actually deploy locally, they're then they're then in charge of their own variable costs. I am then not having to price their variable costs. My view is that the margin you can charge on the variable costs of classification is going to go to $0. So you probably shouldn't be trying to build a business around it anyway. And it sort of makes us happier. It makes them happier. It's way more private. They don't have to fill out paperwork with the Europeans. I don't have to fill out paperwork with the Europeans. It's great. And that really has been a strong attractor for a lot of folks, and is more practical when you're dealing with small fine tune models on top of one or more of the commonly used sort of open weight small models out there because people already know how to run them and can run them on hardware that they've already got set up and configured to run that class of model anyway. Right. So it sort of fits neatly into the various resource planning that folks have to do, which is important because trust and safety teams are usually not first on the draft list for GPUs.

Rob: This is exactly what I was going to ask about, which is obviously you have a startup and so do I, frankly, in the general trust, safety, security area. And how do you think about the people, your advocates inside of these companies are sometimes having to fight for resources. They aren't necessarily the first stop on the resource food chain. So when you're thinking about how you're empowering them to get moved to your solution versus other alternatives that they may have, how do you have that discussion? What do you think in terms of how you help them explain what they need and how they can get it?

Dave: Yeah, it depends where they are. So if they're already employing people to do this, we're way cheaper than that. So that's a fairly straightforward sort of thing to help people understand. It really does depend. We've had a number of customers. We've had some customers that were new enough startups. Runway ML uses us to do a lot of their moderation. And they were building from the ground up and their unit economics don't work if they're having to hire a bunch of people to do this. So they use us to do it. They're able to do actual live moderation of the video because the model is so performant that if you run it on a good enough GPU, you can actually do the classification in real time. So for them, the alternative was impossible, both in the human sense and in the large model sense. Like they're just looking at the numbers and we're like, hey, we fit in there for you. Let's do it. And that worked. There are other folks that had human forces, human workforces that were moving away from them. There are other folks that are maybe a lower resource. So it really depends on their kind of different customer journeys. The interesting thing for me is that I come at this from a policy point of view, obviously, Samidh's background is much more technical and we have customers where our primary context for the policy people who now can basically like roll their own classifiers without really needing engineers. And we have other customers where our primary context at this point are the engineers who can like rely on the robot to write good policy for them, kind of. And that's very interesting. And it feels like a little bit of an echo of the breaking down of expertise silos or like tactical expertise. Right. Because like one of the things the coding models mean, I think, is that the gates are down about the ability to at least make software, not software you should put on the Internet from like a security point of view. But the ability to like tactically make computer do arbitrary thing is now rapidly becoming much more accessible to people who do not have the concrete technocratic skill of writing code. And I think we see in some of our customers the inverse echo of like the ability to do like word stuff sufficiently well coming from an engineering point of view when it's not your core expertise and not something that you really do. And that to me is like a very interesting dynamic that we'll see how it plays out. And then, right, like you end up with resource justification either way, because if you can do enough work for somebody that it saves them one marginal FTE and you cost them less than one marginal FTE, well, like, there you go.

Rob: Yeah, that's super interesting. And I kind of have seen variations of this. Like I've had, you know, companies reach out to me and asked me for advice about stuff where the company was a bunch of engineers, like fast growing startup, and they had no policy people. They're like, I don't know how to build a policy. It's like, well, it's actually, I'm not going to say it's easy, but it's like easier now than it might have been a couple of years ago. And you could look at the competitors, you could do a bunch of things like you can kind of triangulate at least a good starting point, and LLMs can help.

Dave: That's actually interesting, because there's another thing we're seeing, and I'm actually pretty worried about, which is the rise of slop. Yes. In the same way that a lot of slop code, like there's also a lot of bad evaluation. Most of the evals are actually, this is a great point. Evals is a great point to talk about. Yeah, most of the evals are crap. And it's very easy now to get Claude to write a policy, Claude to make examples that do and do not violate it, Claude to grade its homework, and then Claude to tell you it did a good job. And I love Claude's very smart. But that's not a rigorous process. And we all know that's not a rigorous process. But it is a very easy process. And as with code, if you are playing in a domain that is not your expertise, there's a little of a Dunning-Kruger thing that can happen. And so it's interesting, because a lot of our eval period with customers is as much helping them understand how to think about doing evaluation here, as it is selling the product. Like if we get people into an evaluation, and then like hand hold them through like, okay, here's how to do a thorough evaluation, the conclusion is almost always like, oh, this is very good and helpful, and I would like to use it. The trick is getting into an evaluation in the first place, but also making sure that the evaluation is done rigorously. Because if it's not, it's, it's actually quite easy. As a policy-following model gets better at adhering to the policy text you gave it, it can appear worse on badly labeled data than a less highly performing model, which is more insubordinate in the sense that it is ignoring your policy text as literally written and filling in the gaps for you. And so you can get what appears to be a worse result, when actually the problem was just like, both your data and your policy are so shitty that our model is revealing is that your data is not good, and your policy doesn't say what you wanted. And this other model you were using this other approach you were using is just like papering over all of those gaps for you. And so you're not getting what you thought you wanted. And you didn't say what you meant. And so you're not actually in control of your system. But like number is bigger. Right? Because the model is sort of lying to you about whether it's doing what you asked it to do, which has been a very interesting and like nuanced thing to navigate for folks.

Rob: Yeah, that is really interesting. So I mean, does that argue for kind of like, evals as a service?

Dave: Yeah, I think it does. I think it does argue for evals as a service. It's not a place we've gone yet. You know, we've started thinking like, how do we right now, this is a thing we just do with customers. But our first instinct, as an AI forward AI startup is like, can I put together a skill to help them have their agent help them do this themselves? Right? How do we take things we currently do as manual processes for people and wrap up the expertise and a trustworthy, like fair, like concrete way to help that eval go well is definitely a thing that we have batted around as an idea. I don't have the answer to the evaluations yet. But I do think that would be very helpful.

Rob: So where I go with this as well is that, you know, I think of what is AI, what do AI models allow to happen now that couldn't have happened before, just because of availability scale and computing, etc. And so you could even have other entities who have a vested interest in safety, like at a community level at a national level, you know, do some kind of evaluative process from time to time with different companies with different data sets and so on that could again, you couldn't have done this like until very recently. But that might even be something where you have different entities that actually would want to do some kind of spot checks of certain people's or certain companies processes.

Dave: Yeah, we actually have a version of this that has happened. So one of our customers is actually the Meta Oversight Board. They have a case study that they put up on our blog about it and actually also separately released another report where we get name checked in the report. And the one that we did the case study with them on was actually about I believe it was like forced marriage and child marriage content where they're able to do this study over like hundreds of thousands of posts, which like they just would not have been able to do with traditional approaches to labeling because we allowed them to scale up their labeling and their policy tuning. And they've similarly talked about in much more in the vein that you're you're talking about. They've done some work evaluating models responses to various kinds of prompts across I think it's like languages and cultures. And again, used us to help them scale up the evaluating of the vast amount of data they're doing as an organization that, you know, has resources but doesn't have thousands of people to do labeling jobs for them. And it let them do exactly the kind of stuff you're talking about. Now, they're not fully outsiders, right? They're this sort of inside outside hybrid thing. But I think it's a shadow of or an echo of what you're talking about.

Rob: Yeah, I think that's I think that's very interesting. I think also, I mean, one of the things is if liability schemes change in different parts of the world, then you could have a new set of interested actors that, you know, insurance companies and other things that would say, well, you know, the kind of like example I've used before is fire mitigation, right? Your insurance company is not going to cover you if they there's certain things about your property, then if you take care of those things, then they will be like, okay, you're fine. Well, we can ensure you now, that kind of all comes back to what comes back to insurance, I guess, maybe.

Dave: So Airbnb, the secret about Airbnb is it's actually the world's largest multinational insurance company that happens to own a website on which you can get other people's houses. The host guarantee at Airbnb basically means that Airbnb is in the insurance business. So having worked in insurance, I have no desire, but I think you're probably right. And I welcome the insurance companies driving to this her way. Another sort of interesting, like possibility this opens up, which we've seen a little of, but I would love to see more of is actually devolution. Right? So one of the things this means is like, you could put models like this in front of every subreddit moderator, or every Discord space owner, and you could drop them a default set of policies and be like, Hey, guys, change these around. Go nuts, right? But another way of thinking of this is like, if we can miniaturize these tools and make them more accessible and make them have tooling around the tooling to make it easier to access the expertise, you can raise the ceiling of the size of community that an amateur can run. The Reddit situation is really interesting here. My this number is a little bit old. I don't know if it's still true today, but historically, there were about 3000 Reddit moderators who moderated fully 50% of all content on Reddit, because both the unequal distribution of the size of subreddits, and the fact that like moderating a large community is basically a full time job. So these folks who get referred to as power mods, basically, this is just what they do, right? It's become it's become professionalized, even though it's thought of as an amateur thing, at least in the sense of professionalization being how you spend all your time setting aside, right, or relationship to the company. And that is happening because like, when you have a community over a certain size, it is just not feasible for an individual in their free time to manage it to the standard that the central authority requires, without like failing and getting in trouble, basically. And this kind of tooling might allow you to raise that ceiling, and in doing that might make sort of local community on the internet that is more self governing, more possible. This is where I get like to utopian, but I do wonder whether the there was I think, a long standing thesis in civil society that the centralization in social media was due to the lack of data portability, right? And there was this real emphasis on portability for a long time. And if we could break down the silos, then these things wouldn't have to be, you know, so large and monolithic. And we've kind of done a bunch of that, and it hasn't helped. And some of that is just the people are lazy, and it's too late. And there's like network effects. Some of it, I do wonder if it is this, that there are like, this and other kinds of bureaucratic hurdles to managing one of these communities without it being professional. And they're like, maybe these kinds of things make that flattening out more possible, which I tend to think would be good.

Rob: Yeah, those also think there's a there's a drift to, you know, there's the internet, we know the least defended, you know, I guess, you push away abuse from your surface, it goes to less defended surfaces, adjacent surfaces. So that was, water flows down. Yeah, I was reminded of this recently. So I was looking up some scam that I learned about is looking up the name of it. And there was a Trustpilot page for this scam, or mentioning the scam, you know, rating one out of five, like this is a scam, this website is a scam, etc. What was really interesting was some of the top comments on this were essentially saying, yeah, I got scammed out of $10,000. But thank goodness for the scam recovery service that helped me. And it is, you know, the name of it is in the image, my bio image. So, the adversarial nature of some of these problems, like, so basically, like the scam, obviously, crypto recovery scams, there's a bunch of scams that prey on people who have already been scammed. And this essentially was doing that. But they figured out if they put the name of this website in the image, the profile image of the poster on this review site, Trustpilot, in this case, it would get around whatever moderation they have in place. So yeah, no, it will ever be thus, right? It is just a constant conflict forever and ever, as far as I can tell, I try to drop in one random example of something messed up that I've seen in the last couple weeks in every one of these episodes, Dave.

Dave: So yeah, well, I think you're right. Some of the felt urgency in starting this work at Stanford was like, well, the bad guys are obviously going to use the bots to do bad stuff like 10 to a million times faster. So we need to figure out we need to start to figure out how to use LLMs to help defense in the same way, like right the fuck now, because it's gonna take us longer to figure out than it is like figuring out how to defend responsibly and well is slower and harder than figuring out how to destroy things. And so this like embrace of LLMs on the safety side to me, is just like an obvious strategically urgent necessity. And regardless of how you feel about the broader AI spaces, like sort of doesn't matter, like the weapons exist now figure it out, right in the specific context of the adversarial space in which we work, like you have to learn to use these things in all of the ways that we need to learn to use them at a high enough quality immediately. The only better time to start was three years ago.

Rob: One hundred percent. And I think the other thing I just say about this is like, I think, to your point about different folks being able to work on these problems that maybe come from a different perspective, I think there will be alongside the need for these technologies, like the ones you're building at Zentropi, a need for more like community collaboration, discussion about these tools about the adversarial nature of these problems about, you know, is that we're gonna actually, you're gonna also have to, you know, learn the new environment, even if you think you okay, I've been a PM, I've been a policy, but whatever it is, like, we're gonna have to actually learn how to use the new tools effectively together.

Dave: Yeah, I think that's right. And actually, this is like a real wild one that nobody has done. But you could use stuff like what we're building to do. So the policy optimization system we have, we talk about it to people, and most people use it to fix policy text they already have, but you don't actually need to have one, you can just give us a bunch of labeled data, and it'll like, summon a rationalization made out of words for you out of the ether, right? You can start with nothing. And so you could imagine a process where like a community of people took the same examples, and everybody labeled the examples. And then you had this system run and write an explicit rationalization, a policy that is not necessarily the platonically correct one, but is a plausible policy that explains their desires that were expressed just as a bunch of individual choices on individual pieces of content. And then you can take all those documents and throw them into another language model and be like, give me cliff notes on where what we need to go fight about, in order to achieve consensus about how we're going to govern this community. So you could build these like second order tools using stuff like this to produce bottoms up community driven governance in a really interesting way without any of the people involved needing to initially know more than how to rate things against their feelings, which is a pretty soft entry point into this stuff.

Rob: It's the optimized version of organic content policy development: instead of powered by coffee, it's powered by math.

Dave: Yes, it's powered by math. And it's helping smooth the road into a domain that is not actually impossible for normal people to think about. It's just that it has a bunch of technocratic expertise around it. And insofar as we can build reliable tools that have been vetted by people who know what they're doing to help encode some of that technocratic expertise in a system, you then make it more possible for non-experts to enter the space. That's the hope anyway.

Rob: Well, that's great. I think we're out of time. This has been a really fun conversation, Dave. I really appreciate you.

Dave: Yeah, of course. Anytime.

← All Won't Fix episodes