Hard to Vary or Hardly Usable?
Showing only ideas leading to #5562 and its comments.
See full discussion·See most recent related ideasLog in or sign up to participate in this discussion.
With an account, you can revise, criticize, and comment on ideas.My critique of David Deutsch’s The Beginning of Infinity as a programmer. In short, his ‘hard to vary’ criterion at the core of his epistemology is fatally underspecified and impossible to apply.
Deutsch says that one should adopt explanations based on how hard they are to change without impacting their ability to explain what they claim to explain. The hardest-to-change explanation is the best and should be adopted. But he doesn’t say how to figure out which is hardest to change.
A decision-making method is a computational task. He says you haven’t understood a computational task if you can’t program it. He can’t program the steps for finding out how ‘hard to vary’ an explanation is, if only because those steps are underspecified. There are too many open questions.
So by his own yardstick, he hasn’t understood his epistemology.
You will find that and many more criticisms here: https://blog.dennishackethal.com/posts/hard-to-vary-or-hardly-usable
A simple implementation of hard-to-vary
This is my submission for Veritula's bounty for Idea #3069, which asks for an executable implementation that can compare arbitrary English explanations by how hard they are to vary.
Code: hard-to-vary
I've implemented the approach as a small interactive program. It takes arbitrary explanations supplied by the user, collects working variations of each, and ranks the explanations by hardness to vary.
I think hard-to-vary can be implemented more simply than the approaches considered in Dennis's blog, Hard to Vary or Hardly Usable?. The basic idea I'm using is this: if an explanation is harder to vary, there should be fewer ways of changing it while still having it explain what it's supposed to explain.
The program starts by asking the user what question they're trying to answer and then asks for a list of proposed explanations. It goes through each explanation one at a time and asks the user to come up with variations of it that would still work as explanations. The user can enter as many as they can find, then move on to the next explanation.
Finally, the program counts the working variations and ranks the explanations. Fewer working variations means harder to vary, while equal numbers mean equally hard to vary. So if explanation A has two working variations and explanation B has five, A is harder to vary than B.
An example
I'll use Deutsch's example from Chapter 1, “The Reach of Explanations,” of The Beginning of Infinity: Why do seasons occur?
One explanation is the story of Persephone. Hades, god of the underworld, kidnaps Persephone. Her mother Demeter eventually negotiates her release under an arrangement that requires Persephone to return to Hades once a year. Whenever Persephone is away, Demeter becomes sad and makes the world cold and bleak.
Deutsch's point is that we can change many of the details of this explanation while still accounting for the same observations. For example, I might try variations like:
- Persephone escapes instead of being released under an agreement.
- Something other than a magic seed compels her to return.
- Some other arrangement causes Persephone to return annually instead of a marriage contract.
The details have changed, but the story can still be made to explain the seasonal cycle. Suppose I find three such variations.
Now I try the explanation involving Earth's axial tilt. Here the details are much more constrained by what we're trying to explain. Changing the geometry substantially changes what the theory predicts about the seasons. Suppose I only find one working variation.
The program therefore ranks axial tilt as harder to vary: one working variation versus three. That's the entire comparison procedure.
How I understand Dennis's criticism
I don't understand Dennis as claiming that Deutsch simply gets the Persephone example wrong. The issue in the blog is how we get from examples like this, where we seem to have an intuition that one explanation is harder to vary, to a sufficiently specified procedure that could compare explanations in general.
Dennis initially explores numerical quality scores, but that immediately creates problems. Why should one explanation have a score of 500 rather than 550? Why choose that scale? How do criticisms affect the score? How do criticisms of criticisms affect it? I agree with Dennis that these choices look arbitrary.
He eventually gets rid of the quality scores entirely. Instead of trying to measure the quality of an idea, his program keeps track of pending criticisms. His proposed rule becomes: adopt ideas without pending criticisms and reject ideas that have them.
There's something important about how that system works, though. The program doesn't generate criticisms itself; people do. And that's intentional: Dennis says "I’m not looking to formalize or automate creativity as a whole". Creative input can come from users while the program handles the non-creative part of the process. I agree that you can have a rational decision-making process while outsourcing the creative part to the user.
I just don't see why we can't do the same thing with HTV. Let the user come up with variations and say which ones they think still work. The program doesn't need to understand the explanation or come up with the variations itself. It keeps track of the variations and compares the counts.
This doesn't seem fundamentally different from Dennis letting the user tell his program that something is a criticism. In fact, in the blog he explicitly avoids having the program figure out whether a comment is really a criticism: the user checks a box saying that it is. So in my program the user is supplying a working variation; in his, the user is supplying a criticism. In both cases the user is providing the part that requires understanding and judgment, and the program does something simple with that input.
It also means I don't need the quality scores Dennis runs into trouble with. I don't need to decide how many points axial tilt gets compared with Persephone. I'm just asking: how many ways have we actually found to change each explanation while still having it work?
But isn't this subjective?
One criticism Dennis quotes is:
“Also, isn’t the difficulty of changing an explanation at least partly a property not of the explanation itself but of whoever is trying to change it? If I’m having difficulty changing it, maybe that’s because I lack imagination. Or maybe I’m just new to that field and an expert could easily change it.”
I think this is true. I just don't think it's a problem specific to HTV.
Imagine I can't think of any working variations of an explanation, but an expert can immediately think of five. Then yes, our results will be different. But isn't that also what happens with criticism? I might look at an idea and fail to see anything wrong with it while someone who knows much more about the subject immediately sees a serious criticism.
The same goes for participation. An idea on Veritula might have zero pending criticisms simply because hardly anyone has tried to criticize it. That doesn't mean there are literally no criticisms of it. Someone could find one tomorrow. Dennis's answer in the blog is basically that if this bothers you, try to find a criticism yourself. If you can't find one, why not adopt the idea?
I think HTV can work the same way. Zero working variations doesn't mean that we've somehow proven there are no possible variations. It means we haven't found one. If you think the explanation is actually easy to vary, try to come up with a variation that still works.
So yes, the result depends on the knowledge and creativity of the person using the program. But I think rational decision-making is always going to depend on what criticisms, arguments, alternatives, etc. a person is actually aware of. I don't see how Veritula escapes that either.
What about human judgment?
There's another obvious question: who decides whether a variation actually works?
For this program, the user does. I don't think we can get rid of that kind of human judgment, at least until we get AGI. Two people can disagree about whether a variation still explains the thing we're trying to explain. They can also disagree about whether two variations are really different or are basically the same variation stated twice.
Again, I think Veritula has the same underlying issue. People still have to decide whether something really is a criticism, whether a countercriticism actually answers it, whether two criticisms are redundant, and so on. Dennis's program can keep track of the structure, but the structure only means something if those judgments make sense.
Take an extreme case. If a malicious moderator rejects every good criticism of an idea and accepts nonsense countercriticisms, the idea could end up showing 0 pending criticisms. I obviously shouldn't look at the 0 and conclude that the idea is rational to adopt. I'd want to read what happened and decide whether I agree with it.
I don't mean this as a criticism specific to Veritula. I think it's just a limit of this kind of approach. At some point people have to make judgments, and people can disagree about them. Ultimately everyone is their own moderator when it comes to their own decision-making. I have to decide which arguments I accept, which criticisms I think have been answered, and so on.
The same is true with my HTV program. If someone gives me a ranking based on ten supposed working variations, I don't have to accept the ranking blindly. I can look at the ten variations and decide that five don't really work and three others are basically duplicates. My result would then be different.
I'm fine with that. I don't think the goal of either program should be to somehow remove judgment from rational thinking. The program gives us a procedure for what to do with the judgments we've made.
What does the program actually contribute?
Dennis writes:
“We can’t just outsource everything to the user – the app has to do some things or it has no value.”
This was actually the part of the blog that made me think about Veritula itself. What is Veritula doing, and what is it outsourcing?
It outsources the interesting creative part to people. People come up with the ideas. People come up with the criticisms. The user can even tell the program whether something they've written is a criticism by checking a box. The program then keeps track of the structure and tells us how many criticisms are pending.
My program is doing something similar, except with variations. People come up with the explanations and the working variations. The program keeps track of them, counts them, and ranks the explanations.
So I don't think I'm outsourcing everything to the user any more than Veritula is. I'm outsourcing the part that requires creativity and judgment. The actual decision rule is implemented in the program.
For Veritula, that rule ultimately depends on whether there are pending criticisms. For my program, it depends on the number of working variations: fewer working variations means harder to vary.
Where I disagree with the blog
I agree with Dennis that the quality sliders in the blog don't work. I also agree with his decision to let users supply the creative input rather than expecting the program to generate it.
Where I disagree is that I think once we allow this same freedom for HTV, we can construct a similarly simple program for it. The user comes up with explanations and tries to vary them while keeping them working. The program counts the working variations and ranks the explanations.
There are still all the normal problems of human knowledge: maybe I missed a variation, maybe I accepted a bad one, maybe someone else would judge things differently. But those same problems exist when we come up with and judge criticisms.
I don't think either program solves those problems, and I don't think it needs to. That's the part people do.
There are still all the normal problems of human knowledge: maybe I missed a variation, maybe I accepted a bad one, maybe someone else would judge things differently. But those same problems exist when we come up with and judge criticisms.
In this context, there’s a serious issue at the heart of HTV in that not all preference formation involves thinking about variants. Oftentimes, preference formation simply doesn’t take this form. Successful rehabilitation of HTV would be universal in this sense: it would reduce all preference formation to HTV somehow.
I’ve given the example of Pop Tarts in the past. You can ask someone why they’re called that. And they might guess some answers (‘because they “pop” out of the toaster’, ‘because the name is fun’, ‘because the filling tastes tart’, etc.). But ultimately the preferable (and correct) answer will be gotten by just looking it up on the company website, say. Not by seeing which guess is “harder to vary”. I think that should count toward the “normal problems” you mention, but it’s one that needs to be addressed.
➜ hard-to-vary git:(master) python3 htv.pyHard-to-Vary Explanation Comparator-----------------------------------What question are you trying to answer?> Why are they called Pop Tarts?Enter explanations one at a time.Press Enter on an empty line when you're done.Explanation 1: Because they ‘pop’ out of the toaster.Explanation 2: Because the name is fun.Explanation 3: Because the filling tastes tart.Explanation 4:============================================================Explanation:Because they ‘pop’ out of the toaster.Enter variations of this explanation that would still work to explain the same thing.Press Enter on an empty line when you're done.Variation 1:============================================================Explanation:Because the name is fun.Enter variations of this explanation that would still work to explain the same thing.Press Enter on an empty line when you're done.Variation 1:============================================================Explanation:Because the filling tastes tart.Enter variations of this explanation that would still work to explain the same thing.Press Enter on an empty line when you're done.Variation 1:============================================================Question:Why are they called Pop Tarts?============================================================HARDNESS-TO-VARY RANKING============================================================Rank 1Explanation: Because they ‘pop’ out of the toaster.Working variations submitted: 0 variationsRank 1Explanation: Because the name is fun.Working variations submitted: 0 variationsRank 1Explanation: Because the filling tastes tart.Working variations submitted: 0 variationsFewer working variations = harder to vary.
In this case, I couldn’t think of any good-faith variants of any of the three guesses. So they end up being equally preferable according to the program. But I happen to know that all three are false.
If I didn’t know, I’d reject guess 3 because Pop Tarts don’t actually taste tart. They taste sweet. So that’s a pending criticism. But guess 1 and 2 are plausible and I see no reason to prefer one over the other. One would need additional info, but not variant guesses.
I think the Pop-Tarts example mixes together two different things: acquiring new knowledge and forming a preference between explanations. Looking up why Pop-Tarts were actually given their name gives us new knowledge; it doesn't by itself show that the preference formation between explanations isn't HTV.
As I argued in #5553, the program operates on the user's current understanding of the problem and explanations. The user is comparing explanations and variations they currently judge to work. So everything they already know, including criticisms they're aware of, is part of that judgment.
In the Pop-Tarts example, say the three explanations are equally hard to vary given what I currently know. That's fine. I currently have no reason to prefer one over another.
Now suppose I look it up and find historical evidence saying the name was chosen for reason X. I've gained new knowledge, and that changes the problem situation. I'm no longer just asking why they're called Pop-Tarts; I'm asking why they're called Pop-Tarts given that I also know this historical evidence says X.
I could still maintain that the real reason was something else. Maybe the company was lying or the historical source was wrong. But then I have to explain that too. If I just say “the real reason is that they pop out of the toaster, and for some unrelated reason the company says X,” I've added another arbitrary factor to preserve my explanation. That seems like exactly the kind of thing that would expose it as easy to vary.
So I don't see acquiring new knowledge as something that sits outside the process and simply overrides the HTV result. New knowledge changes the problem situation and therefore changes what a working explanation now has to account for. We can then compare the explanations again in light of that new problem situation.
Being hard to vary also doesn't guarantee that an explanation is true. Newton's theory of gravity was a good explanation and remains good enough for many engineering problem situations, even though we later learned that it isn't the full story. As new problems and observations came up that Newton couldn't adequately account for, the problem situation changed and we needed a better explanation.
So I don't think every act of acquiring knowledge or finding a criticism itself needs to take the form of thinking about variants. Those things change the problem situation and affect which explanations and variations we consider to still work.
But that's different from saying the preference formation itself isn't HTV. Once our current knowledge and criticisms are taken into account, HTV can still operate on the explanations that still work and give us a preference—or sometimes tell us that, given what we currently know, they're equally good.
Being hard to vary also doesn't guarantee that an explanation is true.
Agreed, but I’m not asking for a guarantee. I’m basically saying preference formation using HTV should help us find truth. The Pop-Tart example seems to be a case where HTV is unrelated to finding truth.
How does this show that HTV is unrelated to finding truth?
You say:
But guess 1 and 2 are plausible and I see no reason to prefer one over the other.
That's exactly how I see it. Before looking up the historical evidence, both seem like good explanations given what we know. So HTV giving us no preference between them seems like the right result at that point.
Once we acquire new knowledge that distinguishes them, the problem situation changes, as I explained in #5562.
I don't see why HTV failing to distinguish two good explanations before we have the knowledge that distinguishes them means it's unrelated to truth.
So I don't see acquiring new knowledge as something that sits outside the process and simply overrides the HTV result.
One would reject some explanations for reasons unrelated to HTV. It’s not that one would override a specific HTV result – one wouldn’t even get to that result because one would first make other choices leading to different inputs.
If I gather new info and the problem situation changes to where I already know Pop-Tarts aren’t tart but sweet (eg because I taste test them), then I simply won’t include that guess in the inputs to the program, and I won’t bother varying that guess.
HTV could still be part of the process, but DD’s claim was that all rationality boils down to HTV.
One would reject some explanations for reasons unrelated to HTV. It’s not that one would override a specific HTV result – one wouldn’t even get to that result because one would first make other choices leading to different inputs.
I think this is fair. You can treat new knowledge and criticism as changing the problem situation within which HTV operates, or more narrowly as filtering which explanations still work before applying HTV. I think those are functionally equivalent; if you want to call the latter preference formation outside HTV, I'm fine with that.
HTV could still be part of the process, but DD’s claim was that all rationality boils down to HTV.
I'm not defending DD's claim. My claim is that HTV can be implemented.