Search

Ideas that are…

Search Ideas


12 ideas match your query.:

I agree that HTV has to stand on its own. My argument isn't “Veritula has this problem too, therefore it's okay for HTV to have it.”

The reason I've been comparing the two is narrower: to show that outsourcing the parts requiring creativity and judgment to the user doesn't prevent us from having an executable decision procedure built on top of those inputs.

On revisability specifically, I agree that the result should be tentative and able to change when the user's judgments change. As I explained in #5559, this can technically already be done by rerunning the program with the updated inputs. The current CLI is only a minimal demonstration, so that's obviously not ideal UX. Letting the user add, edit, or delete variations and immediately recalculate the result would be straightforward and wouldn't change the decision procedure.

#5560​·​Jad Elmourad, 14 minutes ago​·​Criticism

I think the program already gets this through human input.

The prompt isn't asking the user to enter arbitrary changes. It specifically asks:

“Enter variations of this explanation that would still work to explain the same thing.”

So by entering a variation, the user is making the judgment that it still accounts for what the original explanation was supposed to account for. That's intentionally a human judgment, just like the other judgments I've discussed in #5553.

If the user later changes their mind about one of those judgments, then yes, the current minimal CLI requires them to run it again with the corrected input. I agree that wouldn't be good UX for a full application, but that's not what I'm trying to build here. Adding the ability to add, edit, or delete variations and immediately recalculate the ranking is straightforward and doesn't change the decision procedure being demonstrated.

#5559​·​Jad Elmourad, about 2 hours ago​·​Criticism

I don't think the program needs to choose one of the surviving variants if our current knowledge gives us no reason to choose.

If several variants of the better explanation still work and we currently have nothing that differentiates them, then I think it's fine for them to remain competing possibilities. In fact, my guess is that this is exactly how different research programs can get started: people can follow the different explanatory threads until further criticism, experiments, observations, or other new knowledge gives us a reason to distinguish them.

So I don't think the program should invent a preference between variants when our current problem situation and knowledge don't provide one. Sometimes the rational state really is that we don't yet know which variant is right. We can continue developing and testing them until we find something that differentiates them.

#5558​·​Jad Elmourad, about 2 hours ago​·​Criticism

I agree that people can disagree about what counts as a variation versus a completely different explanation. I don't agree that the program itself needs to recognize the difference.

The program isn't intended to distinguish a variation from a completely different explanation. That's something the user has to decide. In fact, I don't see how a program could make that distinction in general without being a general intelligence that actually comprehends the problem situation and the explanations involved.

And people themselves can disagree about what counts as a variation. That's unavoidable because it depends on how they understand the underlying explanatory argument. What looks like a harmless change to me might look to someone with a deeper understanding like a change that completely breaks the explanation.

So I don't think different people getting different results is necessarily a problem. If I understand an explanation poorly, I may think lots of its details can be changed independently. Once I understand more of the connections between those details, I may realize that many of those variations don't actually work. On the other hand, an expert might also know valid ways of varying an explanation that I would never have thought of. The result reflects our current understanding, and that can change as our knowledge changes.

This is also part of the broader point I made in #5553: the program operates on the user's current understanding of the problem and explanations rather than independently understanding them for the user.

#5557​·​Jad Elmourad, about 2 hours ago​·​Criticism

I think this is where the problem/question given to the program matters.

If the problem I'm trying to answer is “Why does Earth have seasons?”, then replacing Earth with Mars doesn't give me another working variation of the explanation of Earth's seasons. I've changed the phenomenon I'm trying to explain. Applying the same underlying theory to explain seasons on Mars seems to me like an example of its reach, not a variation of its explanation of Earth's seasons.

This is one reason the program asks for the question first. Whether a proposed variation still works can't be judged independently of what we're trying to explain.

As I explained in #5553, the program relies on the user's understanding of both the problem situation and the explanation. So I agree that reach and easy-to-vary shouldn't be confused, but I don't think the program requires us to confuse them. The user has to judge whether they've varied an explanation while preserving what it was supposed to explain, or instead applied the explanation to a different problem.

And of course people can disagree about that distinction, because they can understand the problem situation and explanatory argument differently. That's part of the human judgment the program deliberately leaves to the user.

#5556​·​Jad Elmourad, about 2 hours ago​·​Criticism

As I argued in #5554, if the number of green hats is genuinely part of an explanation and can be changed arbitrarily without affecting anything, then I think that really does expose an easy-to-vary part of the explanation.

But I don't think adding “while I wear one green hat,” “while I wear two green hats,” etc. to the axial-tilt explanation necessarily varies the explanation at all.

If my explanation is that Earth's seasons result from its axial tilt and the resulting relationship between Earth and the Sun, whether I happen to be wearing one green hat or a thousand isn't part of that explanatory argument. Appending an unrelated fact to the English sentence doesn't change the explanation it expresses.

If, on the other hand, I genuinely claim that my hat is part of the explanation—say, that Earth's seasons occur because of axial tilt and because I'm wearing exactly one green hat—then changing the number of hats would be a genuine variation. But in that case I really have made the explanation worse by adding an arbitrary component that isn't constrained by what I'm trying to explain.

So I think there's a difference between varying an explanation and varying the string of English used to express it. The program relies on the user to understand that distinction, for the same reason I described in #5553: it operates on the user's understanding of the explanatory argument, not on arbitrary syntactic changes to its wording.

On that understanding, I don't think both explanations become equally easy to vary merely because infinitely many irrelevant statements can be appended to the sentences expressing them.

#5555​·​Jad Elmourad, about 2 hours ago​·​Criticism

As I explained in #5553, the program operates on the user's current understanding of the explanation.

In this case, though, I think the green-hat example actually illustrates what it means for an explanation to be easy to vary.

Suppose my explanation really is “God causes the seasons while wearing one green hat,” and I genuinely think I can change that to two hats, three hats, a thousand hats, etc. without affecting the explanation at all. Then the number of hats is completely unconstrained by what I'm trying to explain. Nothing in the explanation tells me why it should be one rather than two or a thousand.

I would consider that a problem with the explanation, not with the HTV comparison. It's exactly the kind of arbitrariness that HTV is supposed to expose.

Of course, the program isn't going to enumerate infinitely many hats. It only works with the variations the user actually comes up with. But if I can keep producing arbitrary variations without affecting the explanatory argument, that seems like evidence that the explanation is easy to vary, not a reason to reject the method.

#5554​·​Jad Elmourad, about 2 hours ago​·​Criticism

The program is operating on the user's current understanding of the explanation.
The seasons example in my submission is intentionally a toy example to demonstrate how the program works. In a proper use of the program, both the problem/question and the proposed explanations would include the relevant context needed to understand what is actually being explained and how the explanation works.
For example, “Earth's seasons are caused by axial tilt” is obviously not the full explanation. Someone who properly understands the explanation knows that Earth's tilt matters in relation to the Sun, Earth's orbit, the changing angle and duration of incoming sunlight, the resulting heating, and so on. These parts are connected and constrain each other.
So if the problem is to explain Earth's seasons, I wouldn't accept replacing Earth with Mars, Saturn, Neptune, etc. as working variations of that explanation. Once the relevant context is included, changing the planet changes other parts of the explanatory relationship and no longer explains the original phenomenon.
On the other hand, if a user genuinely understands “because of axial tilt” in such a shallow way that they think Earth can simply be replaced with any other planet while the explanation of Earth's seasons remains intact, then yes: according to their current understanding, the explanation really is easy to vary. The program is correctly reflecting the explanatory knowledge they've supplied to it.
This is why different users can get different results. Someone with a shallow understanding of an explanation may think many of its details can be changed independently. Someone who understands more of the explanatory structure may understand why those same changes break it. Conversely, an expert may also know genuine working variations that a novice would never think of.
So I don't expect the program to produce a context-independent ranking from a few isolated English sentences. The question and explanations are inputs supplied and understood by the user, and the HTV comparison is relative to that current understanding.

#5553​·​Jad Elmourad, about 2 hours ago​·​Criticism

A simple implementation of hard-to-vary

This is my submission for Veritula's bounty for Idea #3069, which asks for an executable implementation that can compare arbitrary English explanations by how hard they are to vary.

Code: hard-to-vary

I've implemented the approach as a small interactive program. It takes arbitrary explanations supplied by the user, collects working variations of each, and ranks the explanations by hardness to vary.

I think hard-to-vary can be implemented more simply than the approaches considered in Dennis's blog, Hard to Vary or Hardly Usable?. The basic idea I'm using is this: if an explanation is harder to vary, there should be fewer ways of changing it while still having it explain what it's supposed to explain.

The program starts by asking the user what question they're trying to answer and then asks for a list of proposed explanations. It goes through each explanation one at a time and asks the user to come up with variations of it that would still work as explanations. The user can enter as many as they can find, then move on to the next explanation.

Finally, the program counts the working variations and ranks the explanations. Fewer working variations means harder to vary, while equal numbers mean equally hard to vary. So if explanation A has two working variations and explanation B has five, A is harder to vary than B.

An example

I'll use Deutsch's example from Chapter 1, “The Reach of Explanations,” of The Beginning of Infinity: Why do seasons occur?

One explanation is the story of Persephone. Hades, god of the underworld, kidnaps Persephone. Her mother Demeter eventually negotiates her release under an arrangement that requires Persephone to return to Hades once a year. Whenever Persephone is away, Demeter becomes sad and makes the world cold and bleak.

Deutsch's point is that we can change many of the details of this explanation while still accounting for the same observations. For example, I might try variations like:

  • Persephone escapes instead of being released under an agreement.
  • Something other than a magic seed compels her to return.
  • Some other arrangement causes Persephone to return annually instead of a marriage contract.

The details have changed, but the story can still be made to explain the seasonal cycle. Suppose I find three such variations.

Now I try the explanation involving Earth's axial tilt. Here the details are much more constrained by what we're trying to explain. Changing the geometry substantially changes what the theory predicts about the seasons. Suppose I only find one working variation.

The program therefore ranks axial tilt as harder to vary: one working variation versus three. That's the entire comparison procedure.

How I understand Dennis's criticism

I don't understand Dennis as claiming that Deutsch simply gets the Persephone example wrong. The issue in the blog is how we get from examples like this, where we seem to have an intuition that one explanation is harder to vary, to a sufficiently specified procedure that could compare explanations in general.

Dennis initially explores numerical quality scores, but that immediately creates problems. Why should one explanation have a score of 500 rather than 550? Why choose that scale? How do criticisms affect the score? How do criticisms of criticisms affect it? I agree with Dennis that these choices look arbitrary.

He eventually gets rid of the quality scores entirely. Instead of trying to measure the quality of an idea, his program keeps track of pending criticisms. His proposed rule becomes: adopt ideas without pending criticisms and reject ideas that have them.

There's something important about how that system works, though. The program doesn't generate criticisms itself; people do. And that's intentional: Dennis says "I’m not looking to formalize or automate creativity as a whole". Creative input can come from users while the program handles the non-creative part of the process. I agree that you can have a rational decision-making process while outsourcing the creative part to the user.

I just don't see why we can't do the same thing with HTV. Let the user come up with variations and say which ones they think still work. The program doesn't need to understand the explanation or come up with the variations itself. It keeps track of the variations and compares the counts.

This doesn't seem fundamentally different from Dennis letting the user tell his program that something is a criticism. In fact, in the blog he explicitly avoids having the program figure out whether a comment is really a criticism: the user checks a box saying that it is. So in my program the user is supplying a working variation; in his, the user is supplying a criticism. In both cases the user is providing the part that requires understanding and judgment, and the program does something simple with that input.

It also means I don't need the quality scores Dennis runs into trouble with. I don't need to decide how many points axial tilt gets compared with Persephone. I'm just asking: how many ways have we actually found to change each explanation while still having it work?

But isn't this subjective?

One criticism Dennis quotes is:

“Also, isn’t the difficulty of changing an explanation at least partly a property not of the explanation itself but of whoever is trying to change it? If I’m having difficulty changing it, maybe that’s because I lack imagination. Or maybe I’m just new to that field and an expert could easily change it.”

I think this is true. I just don't think it's a problem specific to HTV.

Imagine I can't think of any working variations of an explanation, but an expert can immediately think of five. Then yes, our results will be different. But isn't that also what happens with criticism? I might look at an idea and fail to see anything wrong with it while someone who knows much more about the subject immediately sees a serious criticism.

The same goes for participation. An idea on Veritula might have zero pending criticisms simply because hardly anyone has tried to criticize it. That doesn't mean there are literally no criticisms of it. Someone could find one tomorrow. Dennis's answer in the blog is basically that if this bothers you, try to find a criticism yourself. If you can't find one, why not adopt the idea?

I think HTV can work the same way. Zero working variations doesn't mean that we've somehow proven there are no possible variations. It means we haven't found one. If you think the explanation is actually easy to vary, try to come up with a variation that still works.

So yes, the result depends on the knowledge and creativity of the person using the program. But I think rational decision-making is always going to depend on what criticisms, arguments, alternatives, etc. a person is actually aware of. I don't see how Veritula escapes that either.

What about human judgment?

There's another obvious question: who decides whether a variation actually works?

For this program, the user does. I don't think we can get rid of that kind of human judgment, at least until we get AGI. Two people can disagree about whether a variation still explains the thing we're trying to explain. They can also disagree about whether two variations are really different or are basically the same variation stated twice.

Again, I think Veritula has the same underlying issue. People still have to decide whether something really is a criticism, whether a countercriticism actually answers it, whether two criticisms are redundant, and so on. Dennis's program can keep track of the structure, but the structure only means something if those judgments make sense.

Take an extreme case. If a malicious moderator rejects every good criticism of an idea and accepts nonsense countercriticisms, the idea could end up showing 0 pending criticisms. I obviously shouldn't look at the 0 and conclude that the idea is rational to adopt. I'd want to read what happened and decide whether I agree with it.

I don't mean this as a criticism specific to Veritula. I think it's just a limit of this kind of approach. At some point people have to make judgments, and people can disagree about them. Ultimately everyone is their own moderator when it comes to their own decision-making. I have to decide which arguments I accept, which criticisms I think have been answered, and so on.

The same is true with my HTV program. If someone gives me a ranking based on ten supposed working variations, I don't have to accept the ranking blindly. I can look at the ten variations and decide that five don't really work and three others are basically duplicates. My result would then be different.

I'm fine with that. I don't think the goal of either program should be to somehow remove judgment from rational thinking. The program gives us a procedure for what to do with the judgments we've made.

What does the program actually contribute?

Dennis writes:

“We can’t just outsource everything to the user – the app has to do some things or it has no value.”

This was actually the part of the blog that made me think about Veritula itself. What is Veritula doing, and what is it outsourcing?

It outsources the interesting creative part to people. People come up with the ideas. People come up with the criticisms. The user can even tell the program whether something they've written is a criticism by checking a box. The program then keeps track of the structure and tells us how many criticisms are pending.

My program is doing something similar, except with variations. People come up with the explanations and the working variations. The program keeps track of them, counts them, and ranks the explanations.

So I don't think I'm outsourcing everything to the user any more than Veritula is. I'm outsourcing the part that requires creativity and judgment. The actual decision rule is implemented in the program.

For Veritula, that rule ultimately depends on whether there are pending criticisms. For my program, it depends on the number of working variations: fewer working variations means harder to vary.

Where I disagree with the blog

I agree with Dennis that the quality sliders in the blog don't work. I also agree with his decision to let users supply the creative input rather than expecting the program to generate it.

Where I disagree is that I think once we allow this same freedom for HTV, we can construct a similarly simple program for it. The user comes up with explanations and tries to vary them while keeping them working. The program counts the working variations and ranks the explanations.

There are still all the normal problems of human knowledge: maybe I missed a variation, maybe I accepted a bad one, maybe someone else would judge things differently. But those same problems exist when we come up with and judge criticisms.

I don't think either program solves those problems, and I don't think it needs to. That's the part people do.

#5551​·​Jad Elmourad revised about 2 hours ago​·​Original #5523​·​CriticismCriticized3

Thanks for clarifying! The updated #5546 does make it much clearer.

#5550​·​Jad Elmourad, about 4 hours ago

A simple implementation of hard-to-vary

This is my submission for Veritula's bounty for Idea #3069, which asks for an executable implementation that can compare arbitrary English explanations by how hard they are to vary.

Code: hard-to-vary

I've implemented the approach as a small interactive program. It takes arbitrary explanations supplied by the user, collects working variations of each, and ranks the explanations by hardness to vary.

I think hard-to-vary can be implemented more simply than the approaches considered in Dennis's blog, Hard to Vary or Hardly Usable?. The basic idea I'm using is this: if an explanation is harder to vary, there should be fewer ways of changing it while still having it explain what it's supposed to explain.

The program starts by asking the user what question they're trying to answer and then asks for a list of proposed explanations. It goes through each explanation one at a time and asks the user to come up with variations of it that would still work as explanations. The user can enter as many as they can find, then move on to the next explanation.

Finally, the program counts the working variations and ranks the explanations. Fewer working variations means harder to vary, while equal numbers mean equally hard to vary. So if explanation A has two working variations and explanation B has five, A is harder to vary than B.

An example

I'll use Deutsch's example from Chapter 1, “The Reach of Explanations,” of The Beginning of Infinity: Why do seasons occur?

One explanation is the story of Persephone. Hades, god of the underworld, kidnaps Persephone. Her mother Demeter eventually negotiates her release under an arrangement that requires Persephone to return to Hades once a year. Whenever Persephone is away, Demeter becomes sad and makes the world cold and bleak.

Deutsch's point is that we can change many of the details of this explanation while still accounting for the same observations. For example, I might try variations like:

  • Persephone escapes instead of being released under an agreement.
  • Something other than a magic seed compels her to return.
  • Some other arrangement causes Persephone to return annually instead of a marriage contract.

The details have changed, but the story can still be made to explain the seasonal cycle. Suppose I find three such variations.

Now I try the explanation involving Earth's axial tilt. Here the details are much more constrained by what we're trying to explain. Changing the geometry substantially changes what the theory predicts about the seasons. Suppose I only find one working variation.

The program therefore ranks axial tilt as harder to vary: one working variation versus three. That's the entire comparison procedure.

How I understand Dennis's criticism

I don't understand Dennis as claiming that Deutsch simply gets the Persephone example wrong. The issue in the blog is how we get from examples like this, where we seem to have an intuition that one explanation is harder to vary, to a sufficiently specified procedure that could compare explanations in general.

Dennis initially explores numerical quality scores, but that immediately creates problems. Why should one explanation have a score of 500 rather than 550? Why choose that scale? How do criticisms affect the score? How do criticisms of criticisms affect it? I agree with Dennis that these choices look arbitrary.

He eventually gets rid of the quality scores entirely. Instead of trying to measure the quality of an idea, his program keeps track of pending criticisms. His proposed rule becomes: adopt ideas without pending criticisms and reject ideas that have them.

There's something important about how that system works, though. The program doesn't generate criticisms itself; people do. And that's intentional: Dennis says "I’m not looking to formalize or automate creativity as a whole". Creative input can come from users while the program handles the non-creative part of the process. I agree that you can have a rational decision-making process while outsourcing the creative part to the user.

I just don't see why we can't do the same thing with HTV. Let the user come up with variations and say which ones they think still work. The program doesn't need to understand the explanation or come up with the variations itself. It keeps track of the variations and compares the counts.

This doesn't seem fundamentally different from Dennis letting the user tell his program that something is a criticism. In fact, in the blog he explicitly avoids having the program figure out whether a comment is really a criticism: the user checks a box saying that it is. So in my program the user is supplying a working variation; in his, the user is supplying a criticism. In both cases the user is providing the part that requires understanding and judgment, and the program does something simple with that input.

It also means I don't need the quality scores Dennis runs into trouble with. I don't need to decide how many points axial tilt gets compared with Persephone. I'm just asking: how many ways have we actually found to change each explanation while still having it work?

But isn't this subjective?

One criticism Dennis quotes is:

“Also, isn’t the difficulty of changing an explanation at least partly a property not of the explanation itself but of whoever is trying to change it? If I’m having difficulty changing it, maybe that’s because I lack imagination. Or maybe I’m just new to that field and an expert could easily change it.”

I think this is true. I just don't think it's a problem specific to HTV.

Imagine I can't think of any working variations of an explanation, but an expert can immediately think of five. Then yes, our results will be different. But isn't that also what happens with criticism? I might look at an idea and fail to see anything wrong with it while someone who knows much more about the subject immediately sees a serious criticism.

The same goes for participation. An idea on Veritula might have zero pending criticisms simply because hardly anyone has tried to criticize it. That doesn't mean there are literally no criticisms of it. Someone could find one tomorrow. Dennis's answer in the blog is basically that if this bothers you, try to find a criticism yourself. If you can't find one, why not adopt the idea?

I think HTV can work the same way. Zero working variations doesn't mean that we've somehow proven there are no possible variations. It means we haven't found one. If you think the explanation is actually easy to vary, try to come up with a variation that still works.

So yes, the result depends on the knowledge and creativity of the person using the program. But I think rational decision-making is always going to depend on what criticisms, arguments, alternatives, etc. a person is actually aware of. I don't see how Veritula escapes that either.

What about human judgment?

There's another obvious question: who decides whether a variation actually works?

For this program, the user does. I don't think we can get rid of that kind of human judgment, at least until we get AGI. Two people can disagree about whether a variation still explains the thing we're trying to explain. They can also disagree about whether two variations are really different or are basically the same variation stated twice.

Again, I think Veritula has the same underlying issue. People still have to decide whether something really is a criticism, whether a countercriticism actually answers it, whether two criticisms are redundant, and so on. Dennis's program can keep track of the structure, but the structure only means something if those judgments make sense.

Take an extreme case. If a malicious moderator rejects every good criticism of an idea and accepts nonsense countercriticisms, the idea could end up showing 0 pending criticisms. I obviously shouldn't look at the 0 and conclude that the idea is rational to adopt. I'd want to read what happened and decide whether I agree with it.

I don't mean this as a criticism specific to Veritula. I think it's just a limit of this kind of approach. At some point people have to make judgments, and people can disagree about them. Ultimately everyone is their own moderator when it comes to their own decision-making. I have to decide which arguments I accept, which criticisms I think have been answered, and so on.

The same is true with my HTV program. If someone gives me a ranking based on ten supposed working variations, I don't have to accept the ranking blindly. I can look at the ten variations and decide that five don't really work and three others are basically duplicates. My result would then be different.

I'm fine with that. I don't think the goal of either program should be to somehow remove judgment from rational thinking. The program gives us a procedure for what to do with the judgments we've made.

What does the program actually contribute?

Dennis writes:

“We can’t just outsource everything to the user – the app has to do some things or it has no value.”

This was actually the part of the blog that made me think about Veritula itself. What is Veritula doing, and what is it outsourcing?

It outsources the interesting creative part to people. People come up with the ideas. People come up with the criticisms. The user can even tell the program whether something they've written is a criticism by checking a box. The program then keeps track of the structure and tells us how many criticisms are pending.

My program is doing something similar, except with variations. People come up with the explanations and the working variations. The program keeps track of them, counts them, and ranks the explanations.

So I don't think I'm outsourcing everything to the user any more than Veritula is. I'm outsourcing the part that requires creativity and judgment. The actual decision rule is implemented in the program.

For Veritula, that rule ultimately depends on whether there are pending criticisms. For my program, it depends on the number of working variations: fewer working variations means harder to vary.

Where I disagree with the blog

I agree with Dennis that the quality sliders in the blog don't work. I also agree with his decision to let users supply the creative input rather than expecting the program to generate it.

Where I disagree is that I think once we allow this same freedom for HTV, we can construct a similarly simple program for it. The user comes up with explanations and tries to vary them while keeping them working. The program counts the working variations and ranks the explanations.

There are still all the normal problems of human knowledge: maybe I missed a variation, maybe I accepted a bad one, maybe someone else would judge things differently. But those same problems exist when we come up with and judge criticisms.

I don't think either program solves those problems, and I don't think it needs to. That's the part people do.

Update: responding to criticisms

After posting the first version, several criticisms came up that I initially thought about answering individually. After thinking about them, though, I realized that while they raise different issues, they can all be understood within the same framework. So rather than splitting the discussion across several different criticism threads, I'm adding this section and reposting the submission as version 2.

I don't think these criticisms require changing the program or the main argument above. But they do make one part of the argument worth clarifying: the program is operating on the user's current understanding of the explanation.

Dennis showed a case where the program gives the wrong ranking after more variations are entered for the axial-tilt explanation (#5526), and another where arbitrary details like an incrementing number of green hats can generate indefinitely many variations (#5528). This leads to the related problem that both explanations could have infinitely many variations and therefore appear equally hard to vary (#5529). Dennis also pointed out that some variations might represent the reach of a good explanation rather than make it worse (#5530), while Dirk raised the question of what counts as a variation rather than an entirely different explanation (#5531).

I think these concerns can all be looked at within the same framework: the program is operating on the user's current understanding of the explanation.

Take the example of replacing Earth's tilt with the tilt of Mars or another planet from #5526. If someone thinks they can replace Earth with Mars and the explanation still works as an explanation of Earth's seasons, then yes, according to their current understanding the explanation really is easy to vary.

But someone who properly understands the explanation of how axial tilt creates the seasons knows there's a lot more context. The tilt of Earth matters because of Earth's relationship with the Sun, how the amount and angle of incoming sunlight changes, how that affects heating, and so on. All these parts of the explanation are connected. You can't just replace Earth with Mars while keeping everything else the same and still claim to be explaining Earth's seasons.

My example in the program is obviously just a toy version of the explanation. A proper explanation includes all this context, and I think that's where the real hard-to-vary aspect starts to show. As you understand more of the explanation, you understand why certain details have to be the way they are and can't just be swapped out independently.

The same applies to the green-hat example in #5528. If wearing one hat, two hats, three hats, etc. is seriously part of my explanation, and I genuinely think all of those are valid variations that make no difference to what is being explained, then yeah, that's a bad explanation. It contains something completely arbitrary that I can vary however I want. That's exactly what I would expect an easy-to-vary explanation to look like.

This also addresses the infinity issue in #5529. If an explanation really contains something I can vary indefinitely without affecting its ability to explain the phenomenon, then having indefinitely many working variations isn't an accidental problem with the program. It's telling me something about the explanation: it contains an unconstrained part that I can keep changing without consequence.

This is also how I think about the concern in #5530 that some variations represent reach. If we're explaining Earth's seasons and I replace Earth with Mars, I don't think I've found another working variation of the explanation of Earth's seasons. I've applied the underlying explanation to a different problem. If someone understands those as the same explanation exhibiting reach, that's fine too; the important point is that the user has to understand what explanatory claim is being varied and what phenomenon it's supposed to explain.

The program isn't intended to distinguish a variation from a completely different explanation either, as raised in #5531. That's something the user has to decide. In fact, I don't see how a program could make that distinction in general without being a general intelligence that actually comprehends the problem situation and the explanations involved.

And people themselves can disagree about what counts as a variation. That's unavoidable because it depends on how they understand the underlying explanatory argument. What looks like a harmless change to me might look to someone with a deeper understanding like a change that completely breaks the explanation.

So I don't think different people getting different results is necessarily a problem. If I understand an explanation poorly, I may think lots of its details can be changed independently. Once I understand more of the connections between those details, I may realize that many of those variations don't actually work. On the other hand, an expert might also know valid ways of varying an explanation that I would never have thought of. The result reflects our current understanding, and that can change as our knowledge changes.

Dennis also points out that the current CLI doesn't let the user revise a variation after entering it (#5536) and contrasts this with Veritula, where further input can revise the tentative result (#5537). I agree that these judgments need to be revisable. The current program is just a minimal demonstration, so changing your judgment currently means rerunning it. Adding the ability to add, edit, or delete variations and recalculate the ranking is straightforward and doesn't change the decision procedure I'm proposing. If the user's understanding changes, their inputs can change, and the ranking can change with them.

Finally, Dennis asks in #5534 which variant of the better explanation the user should adopt if several variants still work. I think it's totally fine if we don't currently have an answer. If several variants all seem plausible and we have no way of choosing between them, my guess is that this is exactly the kind of situation where scientific research programs can diverge and follow different threads until we eventually find some differentiating factor between the explanations. The program doesn't need to invent a reason to choose between them when we don't currently have one.

#5541​·​Jad Elmourad revised about 16 hours ago​·​Original #5523​·​CriticismCriticized1

A simple implementation of hard-to-vary

This is my submission for Veritula's bounty for Idea #3069, which asks for an executable implementation that can compare arbitrary English explanations by how hard they are to vary.

Code: hard-to-vary

I've implemented the approach as a small interactive program. It takes arbitrary explanations supplied by the user, collects working variations of each, and ranks the explanations by hardness to vary.

I think hard-to-vary can be implemented more simply than the approaches considered in Dennis's blog, Hard to Vary or Hardly Usable?. The basic idea I'm using is this: if an explanation is harder to vary, there should be fewer ways of changing it while still having it explain what it's supposed to explain.

The program starts by asking the user what question they're trying to answer and then asks for a list of proposed explanations. It goes through each explanation one at a time and asks the user to come up with variations of it that would still work as explanations. The user can enter as many as they can find, then move on to the next explanation.

Finally, the program counts the working variations and ranks the explanations. Fewer working variations means harder to vary, while equal numbers mean equally hard to vary. So if explanation A has two working variations and explanation B has five, A is harder to vary than B.

An example

Take the question: Why do seasons occur?

One explanation is the story of Persephone. She spends part of the year in Hades, Demeter grieves while she's gone, and this accounts for the seasonal cycle.

I might try changing the explanation in a few ways:

  • Persephone spends five months in Hades instead of six.
  • It's Demeter's son who is taken rather than her daughter.
  • Zeus orders Persephone to return periodically rather than her return depending on the details involving the pomegranate.

The details have changed, but the basic explanatory story can still do the same job. Suppose I find three such variations.

Now I try the explanation involving Earth's axial tilt. I might be able to change the stated tilt from 23.4 degrees to 23.5 degrees without substantially affecting the explanation. But if I start making large changes to the geometry while trying to explain the same observations, the explanation stops working. Suppose I find one working variation.

The program therefore ranks axial tilt as harder to vary: one working variation versus three. That's the entire comparison procedure.

How I understand Dennis's criticism

I don't understand Dennis as claiming that Deutsch simply gets the Persephone example wrong. The issue in the blog is how we get from examples like this, where we seem to have an intuition that one explanation is harder to vary, to a sufficiently specified procedure that could compare explanations in general.

Dennis initially explores numerical quality scores, but that immediately creates problems. Why should one explanation have a score of 500 rather than 550? Why choose that scale? How do criticisms affect the score? How do criticisms of criticisms affect it? I agree with Dennis that these choices look arbitrary.

He eventually gets rid of the quality scores entirely. Instead of trying to measure the quality of an idea, his program keeps track of pending criticisms. His proposed rule becomes: adopt ideas without pending criticisms and reject ideas that have them.

There's something important about how that system works, though. The program doesn't generate criticisms itself; people do. And that's intentional: Dennis says "I’m not looking to formalize or automate creativity as a whole". Creative input can come from users while the program handles the non-creative part of the process. I agree that you can have a rational decision-making process while outsourcing the creative part to the user.

I just don't see why we can't do the same thing with HTV. Let the user come up with variations and say which ones they think still work. The program doesn't need to understand the explanation or come up with the variations itself. It keeps track of the variations and compares the counts.

This doesn't seem fundamentally different from Dennis letting the user tell his program that something is a criticism. In fact, in the blog he explicitly avoids having the program figure out whether a comment is really a criticism: the user checks a box saying that it is. So in my program the user is supplying a working variation; in his, the user is supplying a criticism. In both cases the user is providing the part that requires understanding and judgment, and the program does something simple with that input.

It also means I don't need the quality scores Dennis runs into trouble with. I don't need to decide how many points axial tilt gets compared with Persephone. I'm just asking: how many ways have we actually found to change each explanation while still having it work?

But isn't this subjective?

One criticism Dennis quotes is:

“Also, isn’t the difficulty of changing an explanation at least partly a property not of the explanation itself but of whoever is trying to change it? If I’m having difficulty changing it, maybe that’s because I lack imagination. Or maybe I’m just new to that field and an expert could easily change it.”

I think this is true. I just don't think it's a problem specific to HTV.

Imagine I can't think of any working variations of an explanation, but an expert can immediately think of five. Then yes, our results will be different. But isn't that also what happens with criticism? I might look at an idea and fail to see anything wrong with it while someone who knows much more about the subject immediately sees a serious criticism.

The same goes for participation. An idea on Veritula might have zero pending criticisms simply because hardly anyone has tried to criticize it. That doesn't mean there are literally no criticisms of it. Someone could find one tomorrow. Dennis's answer in the blog is basically that if this bothers you, try to find a criticism yourself. If you can't find one, why not adopt the idea?

I think HTV can work the same way. Zero working variations doesn't mean that we've somehow proven there are no possible variations. It means we haven't found one. If you think the explanation is actually easy to vary, try to come up with a variation that still works.

So yes, the result depends on the knowledge and creativity of the person using the program. But I think rational decision-making is always going to depend on what criticisms, arguments, alternatives, etc. a person is actually aware of. I don't see how Veritula escapes that either.

What about human judgment?

There's another obvious question: who decides whether a variation actually works?

For this program, the user does. I don't think we can get rid of that kind of human judgment, at least until we get AGI. Two people can disagree about whether a variation still explains the thing we're trying to explain. They can also disagree about whether two variations are really different or are basically the same variation stated twice.

Again, I think Veritula has the same underlying issue. People still have to decide whether something really is a criticism, whether a countercriticism actually answers it, whether two criticisms are redundant, and so on. Dennis's program can keep track of the structure, but the structure only means something if those judgments make sense.

Take an extreme case. If a malicious moderator rejects every good criticism of an idea and accepts nonsense countercriticisms, the idea could end up showing 0 pending criticisms. I obviously shouldn't look at the 0 and conclude that the idea is rational to adopt. I'd want to read what happened and decide whether I agree with it.

I don't mean this as a criticism specific to Veritula. I think it's just a limit of this kind of approach. At some point people have to make judgments, and people can disagree about them. Ultimately everyone is their own moderator when it comes to their own decision-making. I have to decide which arguments I accept, which criticisms I think have been answered, and so on.

The same is true with my HTV program. If someone gives me a ranking based on ten supposed working variations, I don't have to accept the ranking blindly. I can look at the ten variations and decide that five don't really work and three others are basically duplicates. My result would then be different.

I'm fine with that. I don't think the goal of either program should be to somehow remove judgment from rational thinking. The program gives us a procedure for what to do with the judgments we've made.

What does the program actually contribute?

Dennis writes:

“We can’t just outsource everything to the user – the app has to do some things or it has no value.”

This was actually the part of the blog that made me think about Veritula itself. What is Veritula doing, and what is it outsourcing?

It outsources the interesting creative part to people. People come up with the ideas. People come up with the criticisms. The user can even tell the program whether something they've written is a criticism by checking a box. The program then keeps track of the structure and tells us how many criticisms are pending.

My program is doing something similar, except with variations. People come up with the explanations and the working variations. The program keeps track of them, counts them, and ranks the explanations.

So I don't think I'm outsourcing everything to the user any more than Veritula is. I'm outsourcing the part that requires creativity and judgment. The actual decision rule is implemented in the program.

For Veritula, that rule ultimately depends on whether there are pending criticisms. For my program, it depends on the number of working variations: fewer working variations means harder to vary.

Where I disagree with the blog

I agree with Dennis that the quality sliders in the blog don't work. I also agree with his decision to let users supply the creative input rather than expecting the program to generate it.

Where I disagree is that I think once we allow this same freedom for HTV, we can construct a similarly simple program for it. The user comes up with explanations and tries to vary them while keeping them working. The program counts the working variations and ranks the explanations.

There are still all the normal problems of human knowledge: maybe I missed a variation, maybe I accepted a bad one, maybe someone else would judge things differently. But those same problems exist when we come up with and judge criticisms.

I don't think either program solves those problems, and I don't think it needs to. That's the part people do.

#5523​·​Jad Elmourad, 1 day ago​·​CriticismCriticized6