Jacob Goldstein
Jacob Goldstein spent more than a decade as co-host of the Planet Money podcast. He's also the author of the book Money: The True Story of a Made-Up Thing, which the New…
BLOCKERS - New Audiobook from Bestseller Michael Lewis GET THE AUDIOBOOK EXCLUSIVELY HERE
Gain access to ad-free versions of 20+ podcasts from the Pushkin library along with exclusive bonus episodes and other member benefits.
Celebrated novelists, countless students and at least one billionaire have recently been accused of using AI in their writing. The most commonly cited AI detector in these cases is called Pangram.
Max Spero is the co-founder and CEO of Pangram. His problem is this: Can you build tools that reliably detect AI-generated text and images? And a few years from now, will anybody care?
In this episode, Max explains:
Pushkin. I'm Jacob Goldstein, and this is What's Your Problem? My guest today is Max Spiro. He's the co-founder and CEO of a company called Pangram. And Max's problem is this. How do you build a tool that that reliably identifies whether a piece of writing or an image was created by AI. If you've heard about any of the recent controversies over whether something was written by AI, whether it was a prize-winning novel or an op-ed in a major newspaper, there's a good chance that you've actually heard about Pangram, whether you know it or not, because Pangram is the tool people almost always cite in those cases. It is generally viewed as the most reliable, widely available AI identification tool. Max and I talked about a lot of different things in our conversation. We talked about changing norms around when it's OK to use AI and when it's not OK. We talked about the risk of false positives, of PanGram wrongly accusing people of using AI to write things. And we talked about the new PanGram tool to detect whether an image was AI generated. But to start, we talked about two domains where PanGram is already in wide use, education and publishing. Tell me about PanGram and sort of where you are now. Who are your customers? Who's using it?
01:45
Speaker 2
Our customers, a bunch of customers in education, higher ed, and also like high schools, K-12, they want to know whether their students are submitting AI-generated essays. They want a way to teach writing or, like, have any writing-based assignments without having it sort of just devolve into students using AI.
02:07
Speaker 1
It seems like an extremely strong case. Like, you know, writing is a way of thinking hard about things. I buy that it is useful still for students to write. My sense is that students are using AI a lot and, you know, going to great lengths to sort of humanize what the AI says, to sort of write through AI or do what they have to do. Like, what is your sense of how it is going?
02:31
Speaker 2
Yeah, something like pangram here would be more of like, we're adding friction to the process of cheating. And, you know, it's going to make it less enticing if you have to go start with AI bones and rewrite it.
02:43
Speaker 1
Yes, you can't one-shot your assignment.
02:45
Speaker 2
Exactly.
02:45
Speaker 1
Yeah, yeah.
02:46
Speaker 2
And that's helpful in some way. And, you know, you can encourage students to, you know, start from scratch and not use AI at all.
02:54
Speaker 1
But I mean, certainly there are teachers who are requiring students to do their writing in the room. Right? Like, they're just going hard at, you have to sit here in front of me. So what did somebody call it? AI jail or something? You have to sit here in front of me and write an essay.
03:08
Speaker 2
It's kind of a bummer. Like, I feel like we lose a lot here. We lose the in-class instruction. We lose the ability to assign longer written pieces. It's true. When I was in college, I wrote a lot of like eight and 10, 12 page essays.
03:25
Speaker 1
Yeah.
03:25
Speaker 2
Like those, you can't do that in class. That takes a lot of hours.
03:29
Speaker 1
I mean, but that's just the nature of the world now, right? Like, it sort of is what it is.
03:33
Speaker 2
Yeah, and it's a little bit adversarial, as many of our use cases are, where, like, the way that Pangram is used is somebody is trying to get away with AI writing in a way. And so there is this adversarial component.
03:47
Speaker 1
Yes. I mean, it's got to be mostly adversarial, right? The weird non-adversarial one would be, I actually wrote this thing, and I don't want people to think it's AI, so I'm going to put it in Pangram, right? Like, that's an edge case. Like, I've started using fewer em dashes, actually.
04:01
Speaker 2
Oh, man. You got to use the em dashes.
04:04
Speaker 1
Journalists love em dashes. And actually, maybe because I love them, I didn't even notice it in AI pros. But once everybody started talking about the em dash as a giveaway, I just started doing, like, shorter independent clauses, basically.
04:18
Speaker 2
Em dashes are, like, they're no longer an AI tell. This was, like, a GPT-4 era thing. maybe like 4.0, really overuse the em dash. We're over it now. Em dashes are used a normal amount.
04:29
Speaker 1
I can bring the em dash back? I can bring it back?
04:31
Speaker 2
Yes, you should bring it back.
04:33
Speaker 1
Okay, great. So education is obviously a big use case.
04:38
Speaker 2
I'll tell you about publishing as well.
04:40
Speaker 1
Oh yeah, publishing. Well, so that one has been in the news, right? There have been a couple of people, one person lost a seven-figure book deal, right? Because, I don't want to say because of you, but because they were accused of using AI to write the book and you were definitely involved in that.
04:55
Speaker 2
Yeah, there have been a series of scandals this year on AI use, either in books that are already published or like big book deals. But kind of what it comes down to is that a lot of people in the industry are on the left side of the bell curve where they're not heavy AI users. They're like, they're book editors and publishers. They read like real books all day. So they're not like, you know, talking to chat. So I think like a lot of them, they like never really built up this intuition in a way that like all of a sudden, like AI getting good enough to like produce publishable prose was like very surprising and kind of caught them by surprise.
05:39
Speaker 1
So recently, famous investor Stanley Druckenmiller wrote an op-ed in the Wall Street Journal about the bond market. And everybody's like, wait, AI wrote that. And then Druckenmiller was like, you're damn right AI wrote that. I'm 80 years old and I know how to use AI. And then maybe most interestingly, the editor of the journal op-ed page was like, it's fine that AI wrote it. Like Druckenmiller is a famous investor. He really believes this. He stands behind it. Any famous person writing an op-ed in the paper usually has somebody else write it for him anyway. There's no problem here. Sort of surprisingly, I might think I would have a problem with it, but upon reflection, I don't. Do you think it's bad?
06:18
Speaker 2
I think the argument against is that it still does a disservice to the readers. Like, if you read it, it's, like, full of AI prose.
06:27
Speaker 1
So your critique is as a piece of writing. It's bad writing, and that is the criticism.
06:32
Speaker 2
Yeah, and, like, the Wall Street Journal has great editors. I'm sure, like, he could have just, like, written up a draft himself that was, like, not particularly flowery or whatever and then, you know, sent it to his editors and, like, they help him work it into something that's, like, very readable.
06:49
Speaker 1
You could also argue, like, it certainly would be less objectionable if it just said at the bottom, this was written with the assistance of AI, right? Like, that would seem to, like, if you have any criticism, it's like, well, you should tell me because it's weird. Yeah. Although you could make the same argument for ghostwriters and, like, pretty much anything you see written by a CEO, in most cases, of a big company is going to be ghostwritten, right? Like, there are people whose job is writing op-eds for CEOs.
07:17
Speaker 2
I'm going to make this other argument now, which I think is maybe more relevant. So I think there's like two types of AI use and AI writing. There's basically making something more concise or distilling. So for example, if Druckenmiller's like, hey, I wrote like, you know, five pages, here's all my research, here's all my thoughts. I need AI to help me like make this concise and get it into like op-ed size. And so I think typically AI writing in this case ends up being quite helpful because there's a lot to work with and it's distilling it down to its core components. The other side of AI writing that I see is expanding or padding, where somebody's like, here's a few bullet points. Here's a few things that I believe. Turn this into an email or an op-ed. I think that's where it's really doing a disservice to the readers because you're just putting more words that are basically padding and don't mean that much.
08:10
Speaker 1
Just publish the bullet points, bro.
08:12
Speaker 2
Exactly, exactly.
08:13
Speaker 1
Yeah, publish it in Axios. They love bullet points over there. There's a broader thing that's interesting to me about the Druckenmiller Wall Street Journal thing, and that is, like, it does make me wonder to what extent we're right now in a transitional moment where, you know, I could imagine in a year, a few years.
08:36
Speaker 1
In many settings, saying, did a person use AI to write this, would be like today saying, did they use the internet to do the research? Like, probably, and so what are the answers? Does that seem plausible to you?
08:50
Speaker 2
I think we're absolutely in this space or this time period where the norms are completely up in the air and people are so polarized on one side or the other, or some people are like, I'm going to let AI automate my entire online presence. And other people are swearing they'll never use AI ever. And so we have this really wide range of norms. Nobody's really settled on what is acceptable versus unacceptable. I think AI and LLMs, they're very powerful technology. They could do a lot of work for me. They can go do research for me. They can compile this into a spreadsheet for me to read and peruse.
09:31
Speaker 1
They can determine whether people are using AI to write.
09:35
Speaker 2
Yeah. But yeah, I think the answer is we really haven't settled on any norms societally. And the norms that we will settle to are probably that some degree of AI use is okay, but using AI to replace your voice completely, especially in settings where it's going to be read by thousands or millions, is probably less acceptable. Yeah.
09:59
Speaker 1
Well, and disclosure. I do feel like disclosure can solve a lot. It's also interesting to me. You see this certainly on LinkedIn and on Twitter. There are a lot of things that are so obviously AI. And initially, I was like, don't people know? But now I think maybe people know and don't care. Like, I don't know. It does feel... I want to be clear, I don't like AI prose. I'm not speaking of a moral objection, but just as a matter of taste. And then there is a little bit, like, it feels a little rude. It feels a little spammy. Like, I'm reading this thing, and this person didn't even take the time to lightly edit it to get rid of the obvious AI-isms.
10:38
Speaker 2
Yeah, or if you aren't going to take the time to write it yourself, then I'm so much less inclined to read it.
10:45
Speaker 1
So the question I have is, do the people who are posting that know that it is obviously AI and they just don't care? I mean, I guess it varies, but that's a weird one to me still.
10:54
Speaker 2
I think there's a little bit of a bell curve. There's a lot of people who are either new to using AI or they haven't really used that much AI before, and they can't really look at text and immediately grok that it's AI-generated. They can't immediately tell. And so, yeah, obviously you have the people who are just like, don't use AI much. And then you also have the people who like just discovered this technology and they're like, wow, it like writes so well, this is amazing. And I don't think they fully realized that like regular AI users can read this and immediately smell that it's AI generated. Yeah. Then, then you get to the middle of the bell curve where it's like, yeah, like, like people who use AI on a somewhat regular basis, they, they can kind of just tell. And I think you get to the other end where there's people who are like so deep in AI psychosis that, that they also can't tell because like 90% of what they're reading is like AI already. They talk to chatbots more than other people and they just like have lost this ability to differentiate. Uh-huh.
11:51
Speaker 1
So maybe we'll all be at that far end of the curve in a year.
11:55
Speaker 2
I really hope not.
11:55
Speaker 1
Yeah, I suppose I also hope not. So the companies that are making frontier models, Anthropic, OpenAI, arguably Alphabet, Google, are giant companies with wild amounts of money, full of geniuses. You'd think those companies could just do what you're doing. They could just build their own AI detectors. Why don't they?
12:22
Speaker 2
I agree. I don't disagree here. But I also think that this is just not in their top 10 priority list. They're building AGI. They're automating white-collar work.
12:35
Speaker 1
Probably the right answer. I mean, there's an answer, which is like, is it contrary to their incentives to do what you're doing, right? Is it in their interest for people to be able to distinguish AI pros from human pros, or is it contrary to their interest?
12:51
Speaker 2
I think it is contrary to their interest. There was an interesting article that I was exploring a couple of years ago, actually, when OpenAI had watermarking technology, they'd built it, but then they didn't want to release it. There was a lot of internal politics on this, and this was largely because they were afraid it would hurt their core user base, which was at least a plurality of students at this time.
13:15
Speaker 1
They would get caught cheating if there was a watermark. It's basically what you're saying.
13:18
Speaker 2
They would get caught cheating, or they would churn to a different AI service that didn't have a watermark.
13:23
Speaker 1
So you brought up watermarks, so let's talk about watermarks. The EU, as I understand, just passed a law requiring... watermarks on text from language models. Is that true?
13:36
Speaker 2
Yes, that's true. And OpenAI and Anthropic are both working on this. Google already watermarks text outputs.
13:43
Speaker 1
Is that bad for your business?
13:45
Speaker 2
I think it's really good that people are taking this problem seriously. It is competitive with us, but I think Pangram has a lot of advantages that watermarks don't have, which is that a watermark is just going to tell you either Claude was here or not, and Pangram will be able to tell you with higher granularity of like, What was the degree of AI assistance? Was it just helping or did it write the whole thing?
14:05
Speaker 1
I mean, presumably, there's some universe where Claude outputs the copy that's watermarked, and then you literally just retype it into a Google Doc. I mean, naively, I would think the watermark would not follow. Or is the watermark somehow encoded in the prose so that if you retype it, it's still there?
14:25
Speaker 2
Yes, the watermark is statistically encoded in the word choices themselves.
14:30
Speaker 1
Oh, that is rad. So even retyping it, the watermark is still there.
14:35
Speaker 2
Yes. With that said, you could run some automated paraphraser to just paraphrase all the words, and that would erase the watermark.
14:45
Speaker 1
Yes. Well, so there are humanizers now, right? There are things you can upload your language model pros into that claim to evade detection. Although I will say I tried one, I won't name it, but I tried one in preparing for this interview, and you'll be glad to know it didn't work. I had the language model output pros, I ran it through a humanizer, put it into PanGram, and PanGram said 100% AI.
15:09
Speaker 2
Amazing. Yeah, I mean, this is part of the adversarial nature of the business is We are also training on humanizer outputs because we need to make sure that pangram can detect anything that's been run through a humanizer. Otherwise, it would kind of be trivial to evade pangram.
15:27
Speaker 1
Why can't they beat you? I would think you could be at a degree of randomness at the level of word choice and syntax that would defeat you. Why can't they?
15:41
Speaker 2
I think these are kind of like two, there's two competing poles for text when you're humanizing. You want it to be coherent and, you know, retain its logical flow. And then on the other side, you want to just degrade it or perturb it a little bit enough that it is outside of the range of AI detectors. There are humanizers that can beat pangram and you could do this by just like perturbing the text way more just a lot more.
16:10
Speaker 1
Well, sure. I mean, there's a difference between degrading and perturbing, right? Like, you can make it really bad, but that's obviously not useful.
16:19
Speaker 2
Well, I think degrading is just a more extreme version of perturbing.
16:23
Speaker 1
Well, you could improve. You could perturb it a lot and improve it. Oh, yeah, yeah.
16:28
Speaker 2
But these humanizers, they don't have the resources to actually do that, right?
16:31
Speaker 1
Yeah.
16:31
Speaker 2
Most of them are, like, one- to ten-person companies.
16:36
Speaker 1
There's a more fundamental question here that is interesting to me, and I don't really understand the answer, which is, like, why does AI slop all sound the same? You know what I mean? Like, that's the thing that's so annoying. Why can't you just turn up the volume, not on, like, make it bad, but just, like, don't use that same rhythm. Don't do, like, the one sentence. Don't use these annoying words that are AI words. Like, why is that at a technical level?
17:02
Speaker 2
It's really hard to say. I think... So there's a concept in machine learning called mode collapse. And that means like, you know, once an AI model learns what is best or most common, which is the mode, then it will kind of like collapse the edges of its distribution and just output the mode much more often.
17:24
Speaker 1
Yes. The mode being the modal word, the most common word.
17:27
Speaker 2
Yeah, the most common word or phrase or sentence construction. I think that's part of it. And I think also... a lot of these labs, they do post-training. They'll create a reward model and they'll give the language model a reward. And so I can only suspect that there's somebody in here who's saying there's some reward for, quote, like good writing. And that might include, you know, writing with em dashes is probably correlated with good writing in the pre-AI world. Same thing with the like Like, it's not just X but Y sort of metaphors.
18:00
Speaker 1
Yes. That's another one that I have taken out of my own writing, actually. Yeah, yeah.
18:04
Speaker 2
Like, it's a very reasonable construction, but then you feed it into the reward model, and then the AI learns in its prose to turn that up to 11.
18:12
Speaker 1
It likes it too much, yeah.
18:14
Speaker 2
Yeah. So it just overuses what would have normally been considered good writing.
18:22
Speaker 1
We'll be back in just a minute. Okay, let's talk about false positives, right? This is a thing people care about and have written about. So obviously in any test, you can have a false positive or a false negative. A false negative, you're saying this is human generated and actually it's AI. That one people don't get so upset about. The one people are really afraid of is wrongly accused, right? Pangram says it was AI. Actually, it was written by a person. I've seen you say one in 10,000 is the false positive rate. And I've also seen outside studies that, find a very low false positive rate, a similar order of magnitude, one from the University of Chicago I found particularly compelling. But like, I don't know how many writing assignments students have in high school and college, but it's probably order of magnitude of hundreds, right? And so that means that some share of students are going to be wrongly accused, right? If it's 500 writing assignments, then it would be 5% of students will be wrongly accused by Pan Graham of using AI to write a thing that they actually wrote?
19:33
Speaker 2
Yeah, I do think like 1 in 10,000 is beyond the threshold where like you can use this as a tool and generally typically be confident that when it says it's AI, it's correct.
19:45
Speaker 1
It's true. That's definitely true. Like, I mean, I will say just I was actually accused by a teacher in high school of plagiarizing. That's the old school. And I had not plagiarized, for the record. And so that is the old school human version of this. And I'm sure his error rate was higher than 1 in 10,000. But I am sensitive to it. Like, that feeling of being wrongly accused is a terrible feeling. Mm-hmm.
20:08
Speaker 2
Yeah, so I think one thing that pangram helps with actually is it's much better than like the average human intuition. And it's also better than any other AI detector. So because of that, the existence of pangram is, I truly believe this, reducing the number of false accusations for AI use in a scenario like education.
20:29
Speaker 1
Yeah, that's fair. So like notionally, like, of course, teachers are on the lookout for AI one way or the other if they want students to write it.
20:36
Speaker 2
Exactly. And then the other side, I think, which is really important, is that it's not the only source of information. Once Pangram says that something's AI, we can always go deeper. We can either ask the student for drafts. If you use Google Docs, Pangram actually has this Chrome extension where you could see the full revision history, you could see that the document being typed out.
21:00
Speaker 1
And that history is available even without PanGram, right?
21:02
Speaker 2
Exactly.
21:03
Speaker 1
And if you're using Chrome, you can see the version history and go back.
21:06
Speaker 2
Exactly. Yes. Yeah. So that version history is there. But I think PanGram is very important because you don't want to go look at the version history for every single doc. You only want to look into it once you suspect that something was wrong or there are some AIUs there.
21:21
Speaker 1
That's funny. I thought, well, you could automate looking at the version history. And then I was like, oh, but you could also automate a little script that would make it look like you wrote it. It's the adversarial thing. It is amazing. I was reading an article in New York Magazine about students using AI to cheat. And one of the things they pointed out is like how much time students spend cheating, like trying to use AI and not get caught. And like that amount of time you could have written the paper, right? Which is a classic thing, but maybe it's more fun to cheat. Well, yeah.
21:48
Speaker 2
I mean, I think that's part of the like, friction, like adding friction to this cheating process where fewer students are going to do it if it is time consuming to cheat. If it's just one prompt, I think that's where like a lot of students and faculty, like everyone's kind of been frustrated with this problem. It also sucks when your peers are doing one sentence prompts and getting away with it.
22:12
Speaker 1
Right. That's a bad equilibrium. That's a really bad equilibrium. So, okay, so out of every 10,000 times you say AI had anything to do with it, you only get it wrong once.
22:21
Speaker 2
Correct. It's a little bit more of a spectrum in the middle when somebody believes that they had 50% input into the document and AI had 50% input and then Paragram said it's 100% AI. Like, this is sort of the gray area where it's much harder to measure if we're right or not.
22:41
Speaker 1
Uh-huh. And so that's the area where you get it wrong more is kind of at that margin. So it's like, no, no, I wrote it and I got some edits from AI, but PanGram returns 100% AI. Like that happens more often than 1 in 10,000.
22:56
Speaker 2
Yeah, that happens at least for like light AI assistants or medium AI assistants. It's like a few percent of the time where we say it's still like 100% AI.
23:08
Speaker 1
So a few percent is orders of magnitude higher than one in 10,000. It happens kind of a lot.
23:15
Speaker 2
We try to be very transparent about all of our benchmarks. I think this is also the area of greatest debate where it's not actually so black and white. What the degree of AI assistance was, people on average tend to under-disclose. They're like, oh, I use AI to help me. I talked to someone who's like, I used AI for copy edits. And then you go through what that means. It's like they actually, like, every word was produced by AI. And it had gone through so many revisions with AI. And so I think this is the biggest gray area where Pangram is not going to be, like, a perfect oracle and tell you this was the exact prompt and these were the exact human edits. But I think it can, like, get approximately there.
24:01
Speaker 1
Are there... groups of people or kinds of speakers where pangram is worse at detecting whether their writing is their own or generated by AI?
24:11
Speaker 2
So we've run this on three different groups. I think there's been different studies in the past that have found that AI detectors are biased against English as second language speakers. There's also been questions of if AI detectors are biased against autistic or neurodivergent speakers. And so what we've done is we ran studies where we looked at 25,000 texts by English as second language speakers and found that there's no elevated false positive rate. We have one false positive in the 25,000. And then so similarly, somebody else who is not us asked for credits to run a similar study for autistic speakers and they did a really great job. They had like a 20 page report on their findings and they also found no bias against autistic speakers, which is really great. And then the third one that we've looked at was just like race. Cause I think this has also been a question and we again found no statistically significant difference, which it was really cool to see. What?
25:28
Speaker 1
Yes.
25:29
Speaker 2
Yes.
25:29
Speaker 1
One hopes that would not be a problem. You could imagine how for someone who was not a native speaker, for example, you might have a higher rate of false positives. Was it a problem in the past? Like were earlier versions of PanGram less reliable in that way?
25:43
Speaker 2
I would say like PanGram 2, like definitely had a bit of this issue. And then starting like PanGram 3, we were taking it very seriously. But I think every, we've had a lot of candidate models that were biased against English second language speakers and And so because of that, we kind of threw away the model and never released it.
26:02
Speaker 1
And when was Pangram 2?
26:03
Speaker 2
How long ago was that? Like a year and a half ago. Even Pangram 2, like our false positive rate was like much lower than the other AI detectors. But I think it was still like compared to our baseline false positive rate, it was elevated on ESL. And I think like part of the difficulty here is going back to the humanizers. a lot of them are trying to inject spelling mistakes, grammar mistakes as a way to evade AI detection. And so training on these naively will make your model correlate these spelling and grammar mistakes with AI, actually, which is not correct.
26:45
Speaker 1
Well, only if those are occurring at a higher rate in the humanizers than they are in the human texts you're putting in, right? Yeah.
26:53
Speaker 2
I think the humanizers, some of them are perturbing and degrading the text more than your average human. So it's taking it like all the way down to the other side. Yeah.
27:03
Speaker 1
There have been a number of cases where, you know, Pangram found a piece of writing to be AI generated. The person said, no, it wasn't. I wrote it. Has there been a case where the person was validated? A high profile case?
27:15
Speaker 2
Yeah. So there was like one case, Taylor Lorenz had this like article in Vanity Fair. So Pangram has a 50 word minimum. Somebody selected a 54 word segment and Pangram said this was AI generated. But what they did was they just posted the like a hundred percent AI for the 54 word segment. And so then I looked into it. If I checked like the full Vandy fair article, Pangram actually said a hundred percent human. And then, you know, Taylor's like, Hey, like here's my edit history. Like we went through a lot of provisions with, with an editor, like this is all my writing. And so, yeah, like it's, it was absolutely a false positive and, Yeah, partially because it was a small piece of text and partially, I think, because this person was kind of using pangram in an inappropriate way of, like, I don't know, trying to go for the gotcha rather than, like, going for the full story.
28:06
Speaker 1
Well, that 1 in 10,000 false positive rate, is that contingent on some minimum number of words? And if so, what number of words?
28:14
Speaker 2
This is aggregate over positive. all of our texts, but it's a little bit higher for 50 words.
28:19
Speaker 1
If you put in 50 words, is the false positive rate notably higher?
28:23
Speaker 2
Yes, it's closer to 1 in 5,000. Uh-huh.
28:27
Speaker 1
I mean, maybe the person looked for the most AI-sounding 50.
28:30
Speaker 2
Words, right? I mean, that's what they said, is they said this sounded like AI to them. So it's a false positive to them and also to PanGram.
28:37
Speaker 1
Okay, let's talk about images. Images are new. I think when we booked you, maybe, I don't even know if the images thing was out, but tell me about your image business. Yeah.
28:46
Speaker 2
So it's a new model. We released it in July, end of July, but it is state-of-the-art for AI image detection. So what that means is you can put in an image and we'll tell you, yes, this is AI generated and it works on images from nano banana, like Gemini, GPT image, mid journey, all of the major AI image models. And so I think this is pretty cool because it A couple of years ago, it was like really easy to tell whether an image was AI generated. You just like count the fingers and you look at the text.
29:19
Speaker 1
Count the fingers, right. Or see if there's something that defies the laws of physics, right?
29:23
Speaker 2
Yeah. And today that happens a little bit, but like you can get some really hyper-realistic AI images. It's scary. And I think it's really hard to tell.
29:33
Speaker 1
Certainly hard for people to tell. I mean, I certainly buy that you are at the forefront of language detection. What is the competitive landscape like for image detection?
29:43
Speaker 2
There's a number of AI image detectors. I think many of them are more focused on deepfakes and less focused on AI images in general.
29:54
Speaker 1
Yes, I talked to the head of Lodi. And they're focused in particular on videos. Like his claim, at least, was that they are good at understanding video, ingesting video in a way that other people are not.
30:06
Speaker 2
Oh, that's pretty cool.
30:06
Speaker 1
Because video is hard.
30:08
Speaker 2
Yeah. Yeah, we've been working a bit with video. So if you put in a single frame from a video into the AI image detector, it will do pretty well.
30:15
Speaker 1
What kinds of businesses are natural customers for your image product?
30:19
Speaker 2
Yeah, I think naturally it would be someone like Pinterest or like any like stock photo company. image website, things like that, where you want your users to have some degree of trust that what they're seeing is real.
30:35
Speaker 1
Dating apps or websites was the first one I thought of.
30:38
Speaker 2
Yeah, I think dating apps, absolutely. You don't want AI-generated people or AI catfishing.
30:45
Speaker 1
Real estate apps was another one I thought. I mean, I know they will explicitly sometimes be like, and here's what it could look like. But like, they don't always say that. I feel like a lot of the time, like my neighborhood has all these old houses that people have lived in for 50 years, but they sell for a lot of money now, right? Which is the classic thing where the realtor wants to be like, look, you can pay a lot of money for this house because it could look like this. But they don't show you what it looks like now for a reason.
31:08
Speaker 2
Oh, that's horrendous.
31:09
Speaker 1
Yeah, yeah.
31:10
Speaker 2
I mean, I think that sort of stuff, like it should be illegal to... to put an AI photo instead of the actual photo in your listing?
31:18
Speaker 1
They may well say it. I don't want to be accusing people of malice, but, like... No, I've.
31:23
Speaker 2
Definitely seen it for apartment listings in New York. Yeah. Like, where I show up, and it's like, huh, it looks a little bit different than the photos. And I go back to the photos, and, like, some of the details are different. And I'm like, okay, I think... I'm, like, pretty sure this was AI.
31:36
Speaker 1
Wait, it's a shady apartment listing in New York? No, I don't believe you. I.
31:42
Speaker 2
Whenever it's too good to be true, half the time it's AI.
31:46
Speaker 1
One time I called for an apartment in New York to go look at it, and the guy's first question was, how tall are you? I was like, I'm out. I'm not going to go look. We'll be back in a minute with The Lightning Round. Okay, let's finish with the lightning round. Cool. How has building PanGram changed the way you think about language?
32:22
Speaker 2
I don't know if I have an answer for this.
32:25
Speaker 1
Maybe it hasn't. Maybe it hasn't.
32:26
Speaker 2
No, I mean, I think some of the biggest learnings for me were that, oh, like, these language models are not at all as diverse as I thought. Yeah. When I started, like, ChatGPT had come out, and we were all like, this is amazing. incredible technology. And then spending a little more time with it made me realize, oh, actually, it's kind of just like saying the same thing every time. It's quite mode collapse. And humans are much more diverse. And I think that's valuable.
32:56
Speaker 1
Well, what's the last thing you read that you loved?
32:58
Speaker 2
I read this book called Sublimation by Isabel J. Kim. It was about sort of this concept where if somebody immigrates, they split into two people. And one person stays, one person goes. It was very interesting. It's a good book.
33:13
Speaker 1
It's a novel?
33:14
Speaker 2
Novel, yeah.
33:15
Speaker 1
You ever read anything by AI that you loved?
33:18
Speaker 2
There's this kind of funny green text from an old AI model about being a bottomless pit supervisor.
33:28
Speaker 1
Wait, I don't know what green text is. What's green text?
33:31
Speaker 2
Green text, it's this format of storytelling from 4chan where it's basically describing a story in first person. Every line starts with this carrot or this like.
33:45
Speaker 1
I see. Kind of like looking at a command line or something.
33:48
Speaker 2
Yeah, like the command line thing. And then I think because of that, like the formatting, it was green.
33:53
Speaker 1
Okay, so go on. So you saw something that was like this and it was about?
33:57
Speaker 2
Yeah, it was kind of just like a joke story about being a bottomless pit supervisor. And I was like, okay, this is like, this is kind of funny. This is the first time I've like laughed at something that was AI generated.
34:08
Speaker 1
How did you come across it?
34:10
Speaker 2
It was posted on Twitter. I can read it for you if you like.
34:14
Speaker 1
Yeah, that'd be great.
34:14
Speaker 2
Okay. So it starts with the human prompt is write me a 4chan green text, be me bottomless pit supervisor. And then AI continues this with in charge of making sure the bottomless pit is in fact bottomless. Occasionally have to go down there and check if the bottomless pit is still bottomless. One day I go down there and the bottomless pit is no longer bottomless and The bottom of the bottomless pit is now just a regular pit. Distress.jpg. Ask my boss what to do. He says, just make it bottomless again. I say, how? He says, I don't know. You're the supervisor. Rage.jpg. Quit my job. Become a regular pit supervisor. First day on the job, go down to the new hole. It's bottomless.
35:02
Speaker 1
That's great. It also doesn't sound like AI. Yeah. To me, to my ear, anyway.
35:10
Speaker 2
Yeah, I think that's because this was, like, old AI. I want to say it's, like, GPT-2 or 3.
35:16
Speaker 1
I also wonder, like, it sounds like it's not a form that I'm familiar with, but it sounds like it is this very particular form of writing that exists in the world. And so, just like if you say, write a sonnet or whatever, I don't know, maybe Pangram can tell an AI sonnet from a human-generated sonnet, or maybe it's not enough words. But, like... When there is a form, maybe we all sort of collapse. Humans collapse to the form as well. And so the collapse of AI is maybe less obvious. I don't know.
35:42
Speaker 2
I think so. Yeah. The collapse of the form and adherence to this structure is, I think, makes it interesting. And it's very interesting that AI is able to adhere to the structure so well.
35:54
Speaker 1
And it gets the joke, right? Like bottomless pit supervisor is an absurd premise, right? And it treats it just right. Like it's a deaf comic touch, right?
36:05
Speaker 2
Yeah, it was good.
36:06
Speaker 1
Tell me about your Twitter game, your X game. You're on a lot and you're kind of a poster. Like your persona on Twitter is a little bit more, I don't know if combative is too strong of a word, than your persona in this interview.
36:19
Speaker 2
I like to be on Twitter a lot. I post all the time. I kind of just like post what I think. I don't think too hard about it. But I think Twitter is a really great community because... First off, I like text. I'm not on TikTok. I'm not filming videos of myself. I like just typing things out. So I think it's a great place for me. And it's also kind of like where all the AI happenings happen. So a lot of the people I follow are people who work in AI in various capacities, either AI researchers or independent people doing cool things with AI. So I post a lot.
36:56
Speaker 1
Are you not more combative on Twitter? Am I misreading that?
37:00
Speaker 2
Sometimes I think, so I often will get like, somebody will say, you know, pangram was wrong on like so-and-so and then I'll have to jump in and be like, Hey, you know, can you send me the text? Like we want to go like triage it and figure that out.
37:13
Speaker 1
Yeah.
37:14
Speaker 2
Uh, in terms of combativeness, I, I'm not like picking fights all the time, but I, I mean, I think maybe it's just a function of the platform and the format where like, you know, here we're having such a like nice, pleasant conversation. But, like, there's no pleasant conversations on Twitter.
37:32
Speaker 1
Yeah, it's interesting. I will occasionally get angry emails, as one does. And almost every time, if I write back in a polite way, the person apologizes. I don't even ask them for an apology. I just say, yes, thank you for pointing that out. Here is my thinking. And more often than not, if I do that, the person will write back apologetically, basically, or certainly gratefully. What's your least favorite AI thing?
38:01
Speaker 2
My least favorite is when it's like, and honestly. I'm just like, that's horrible.
38:09
Speaker 1
Load-bearing and doing the work are the two.
38:12
Speaker 2
That, yeah, those make me cringe.
38:23
Speaker 1
Max Biro is the co-founder and CEO of Pangram. Today's show was produced by Gabriel Hunter Chang. It was engineered by Sarah Bruguere. And our showrunner and editor is Jake Harper. I'm Jacob Goldstein, and we'll be back next week with another episode of What's Your Problem?
Jacob Goldstein spent more than a decade as co-host of the Planet Money podcast. He's also the author of the book Money: The True Story of a Made-Up Thing, which the New…