Activity

  Dennis Hackethal criticized idea #5607.

Just to confirm, and so you can hear it from me, yes I did use AI to help with my submission.

For posterity, here was my position during the discussion:

I think it's technically fair for Dennis to consider the submission ineligible because the bounty terms clearly say "Submissions must be handwritten. No AI-generated text or code; suspected AI use makes submissions ineligible."

On the other hand, I think this outcome is a shame because I don't think my use of AI changes any of the real work and thought that was put into it. If you zoom out, the motive of the bounty is to attract people to work on ideas and develop them, which is exactly what I did here, regardless of the tools used. As How Do Bounties Work? puts it, "Bounties let you invite criticism and reward high-quality contributions with real money."

And this is exactly what happened here. The idea had been posted on Veritula for about 10 months, and when I saw the bounty, I took it as a fun challenge to work on. So in that sense, the bounty incentive worked. And it seems to me that the reception of the idea has been good so far, which is why I find this outcome disappointing. On a symbolic level—and I acknowledge that this may be very different from how Dennis sees or intends the decision—his withdrawal of all his funds feels to me like this work is somehow not something he's willing to reward.

I did attempt to persuade Dennis to reinterpret the no-AI rule by discussing the underlying reasons for the rule. We agreed that handwriting can be a way to test the submitter's own understanding rather than relying on AI's understanding. While we agreed that this wasn't an issue for me, we disagreed about what should follow from that. My view was that the purpose of the rule should matter more than the literal wording, whereas Dennis seemed to think that the written terms should still govern.

Ultimately, we couldn't come to an agreement on whether my use of AI was appropriate in the context of this bounty. We also disagreed about whether I should have raised my use of AI proactively. In any case, I respect Dennis's decision since ultimately these were the terms of his bounty. We also agreed that in the future, discussions like this about the terms ought to be brought up at the beginning rather than after the fact.

#5607​·​Jad Elmourad, 1 day ago

When thousands of dollars are at stake, I think it’s fair to ask people to play by the rules. Still, I’m allowing flexibility by giving Jad a chance to collect the remaining funds.

I’m always happy to reward good work; I’ve paid out bounties in the past. And before I found out about Jad’s use of AI, I gave him tips increasing his odds of beating the bounty. I also submitted criticisms of some of my own ideas, further increasing his odds.

But people need to earn the reward. At some point, I noticed Jad had submitted his replies very quickly, with unusually verbose text, and in fairly rapid succession. AI is notoriously verbose, so that would have been the second-best time for him to mention he was using AI. He did not. But some people do genuinely type fast, so I gave him the benefit of the doubt.

When I later asked him directly whether he was using AI, he admitted to it but claimed it was merely to “help edit some of [his] writings”.

Pangram, an AI detector, paints a different picture: it says 100% (!) of the code and 22% of the prose in his first idea are AI-generated. I’ve also sampled some of his replies since, which Pangram says are also 100% AI-generated.

Pangram result: 100% AI-generated

To play devil’s advocate once again, I wanted to make sure Pangram wasn’t giving me false positives. So I ran several tests against some of my own prose and code, and compared the results with known AI-generated prose and code. Pangram was almost always right – the few times it did make mistakes, they were almost always false negatives, so if anything, Pangram is too permissive. It does have lower confidence for shorter texts, so I also made sure to use only texts that are long enough. I showed Jad these results privately and he didn’t have a good answer, so I’m surprised he now wants to continue the discussion publicly.

In terms of content, handwriting the code matters because of Deutsch’s yardstick for having understood a computational task: if you can’t program it, you haven’t understood it. Jad says he still understands the code regardless, so it shouldn’t matter. But why not raise that before agreeing to a rule he took issue with? Why accept the risk of disqualification instead?

In terms of fairness, as Jad knows, payout for a bounty depends on addressing all known criticisms inside a review deadline. When a participant uses AI to help with that, but the funders don’t know about it, it creates an unfair advantage. That’s a violation of not just the letter but also the spirit of the rules. It’d be like secretly using a chess computer for help in a chess competition. The amount of text Jad has submitted also creates an uphill battle for the remaining funders to review his submissions before the deadline. Jad says he knows how bounties work, so he knows payouts are automatic at the deadline. Still, the remaining funders are willing to accommodate him.

Before I made the decision to withdraw my funds, I reached out to Jad privately, as I said, to make sure there wasn’t anything I had missed and that there were no hard feelings. We discussed the issue thoroughly and I couldn’t find anything. I also discussed the matter with the co-funders for a second and third opinion, again to make sure I hadn’t left any stones unturned. (To be clear, the responsibility for the decision to withdraw my funds is mine alone.)

I understand Jad was hoping for a bigger payout, so this must be a disappointing outcome. I empathize with that. I’d be bummed too. But he could have easily avoided this outcome. One can’t agree to a rule, immediately break it, get an unfair advantage, hope I won’t notice, and then complain when I do. But with any luck, he’ll get the remaining funds, provided that the arguing ends now.