Search Ideas
2700 ideas match your query.:
So I don't see acquiring new knowledge as something that sits outside the process and simply overrides the HTV result.
One would reject some explanations for reasons unrelated to HTV. It’s not that one would override a specific HTV result – one wouldn’t even get to that result because one would first make other choices leading to different inputs.
If I gather new info and the problem situation changes to where I already know Pop-Tarts aren’t tart but sweet (eg because I taste test them), then I simply won’t include that guess in the inputs to the program, and I won’t bother varying that guess.
HTV could still be part of the process, but DD’s claim was that all rationality boils down to HTV.
Being hard to vary also doesn't guarantee that an explanation is true.
Agreed, but I’m not asking for a guarantee. I’m basically saying preference formation using HTV should help us find truth. The Pop-Tart example seems to be a case where HTV is unrelated to finding truth.
@jadelmourad be sure to submit separate criticisms separately. You benefit because it means I need to address each criticism individually, not in ‘bulk’. (See #5546, keyword ‘bulk’.)
One could run the program again with corrected inputs but as I noted in #5533, that approach won’t scale.
Jad originally replied in #5559. I’ve extracted his reply to avoid bulk criticism:
If the user later changes their mind about one of those judgments, then yes, the current minimal CLI requires them to run it again with the corrected input. I agree that wouldn't be good UX for a full application, but that's not what I'm trying to build here. Adding the ability to add, edit, or delete variations and immediately recalculate the ranking is straightforward and doesn't change the decision procedure being demonstrated.
I think the program already gets this through human input.
The prompt isn't asking the user to enter arbitrary changes. It specifically asks:
“Enter variations of this explanation that would still work to explain the same thing.”
So by entering a variation, the user is making the judgment that it still accounts for what the original explanation was supposed to account for. That's intentionally a human judgment, just like the other judgments I've discussed in #5553.
I don't think the program needs to choose one of the surviving variants if our current knowledge gives us no reason to choose.
It does. Preference formation is the point of the program.
Sometimes the rational state really is that we don't yet know which variant is right. We can continue developing and testing them until we find something that differentiates them.
Then we’d need to find a way to feed the results of the tests back into the program. Tests would somehow need to be formally tied to HTV. The program has to at least be able to tell the user which preference to form (short of the user not typing in the variant he’s already decided he doesn’t want, eg because of a negative test case).
But suppose grass really did cure the disease and we actually understood how. Maybe a particular compound in the grass interacts with some biological mechanism, and a certain concentration is required for the effect. The explanation might allow a whole range of effective dosages rather than one exact number, but that range would itself be explained and constrained by the mechanism.
That assumes the cure works: then I agree there’s a reason for the specific range, and that range will be part of the one working variant. Your program would then accurately prefer it over a rival.
But the program also needs to account for cases where the cure doesn’t work. Then we can’t refer to a range anymore because there’s no reason to constrain the explanation to that range. And then we’re left with the situation described in #5549.
I don’t think this response addresses the part from #5531 that says “nor is the user given an opportunity to correct any mistakes to that effect.”
So I don't think different people getting different results is necessarily a problem.
#5531 doesn’t say that.
But I don't think adding “while I wear one green hat,” “while I wear two green hats,” etc. to the axial-tilt explanation necessarily varies the explanation at all.
If you don’t think those are genuine variations, the program should either reject them or, in light of #5553, it should give users a way to resolve disagreement about what makes a genuine variation, because such a resolution would presumably itself involve some measure of HTV and so is part of the problem of making HTV work.
As the bounty states, human input is fine within reason. It’s beginning to look like your program does relatively little compared to the amount of work it expects its users to do.
I would consider that a problem with the explanation, not with the HTV comparison. It's exactly the kind of arbitrariness that HTV is supposed to expose.
Yes but if the program is an accurate representation/use of HTV, then the program should expose it. The program asks the user for variations in service of exposing it. But that doesn’t work in this case. I think that’s a problem.
As I recall, FoR has an example of a grass cure. It says something like: you could eat 1g of grass to cure some disease. If that doesn’t work, proponents of that theory can simply say you actually need to eat 0.9g. When that doesn’t work either, they can say you need 0.95g. And so on. So there is literally an uncountably infinite number of variant theories: one for each decimal number in that vicinity.
Put those variants up against a competing explanation that also has uncountably many variants. Now you can arbitrarily get a result saying to prefer one over the other by typing in fewer variants for one than for the other.
More abstractly put, we can’t always count the variants. But the algorithm, as written, expects us to.
Revisions aren’t for counter-critizing criticisms. Revisions are for changes you make because you agree with a criticism. I’ve restored the original. I’ve also updated #5519 to make the purpose of revisions clearer. It’s now #5546. Please read it in its entirety before you continue working on the bounty. Then continue with #5544.
How Does Veritula Work?
Veritula (Latin for ‘a bit of truth’) can help you live a life guided exclusively by reason.
To reason, within any well-defined epistemology, means to follow and apply that epistemology. Unreason, or whim, is an undue departure from it. Epistemology is the study of knowledge – basically, the study of what helps knowledge grow, what hinders its growth, and related questions.
Veritula follows, and helps you apply, Karl Popper’s epistemology, Critical Rationalism. It’s a continuation of the Athenian tradition of criticism and the only known epistemology without major flaws.1
Critical Rationalism says that ideas are assumed true until refuted. This approach leaves us free to make bold guesses and use the full arsenal at our disposal to criticize these guesses in order to solve problems, correct errors, and seek truth. It’s a creative and critical approach. Critical Rationalism is a fallibilist philosophy: there is no general criterion of truth to determine with certainty whether any given idea is true or false (Tarski). We all make mistakes, and by an effort, we can correct them to get a little closer to the truth. Rejecting all forms of mysticism and the supernatural, Veritula recognizes that progress is both possible and desirable, and that rational means are the only way to make ongoing progress.
Veritula is a programmatic implementation of Popper’s epistemology.
Veritula provides an objective, partly automated way to tentatively determine whether a given idea is problematic. It does not tell you what to think – it teaches you how to think.
On Veritula, ideas are discrete and immutable. Consider an idea I:
I
Since it has no criticisms, we tentatively consider I unproblematic. It is rational to adopt it and act in accordance with it. Conversely, it would be irrational to reject it, consider it problematic, or act counter to it. (See this quick guide for more details on rational decision-making.)
Next, someone submits a criticism C1:
I|C1
The idea I is now considered problematic so long as criticism C1 is not addressed. How do you address it? If you agree with the criticism, you can revise I so that C1 doesn’t apply anymore, which restores the previous state with just the standalone I (now called I2 to indicate the revision):
ReviseI ------------> I2|C1
To track changes, Veritula offers beautiful diffing and version control for ideas.
If you disagree with C, you can counter-criticize it instead, thereby neutralizing it with a new criticism, C2:
I|C1|C2
Now, I is considered unproblematic again, since C1 is problematic and thus can’t be a decisive criticism anymore.
If you can think of neither a revision of I nor counter-criticism to C1, your only option is to accept that I has been (tentatively) defeated. You should therefore abandon it, which means: stop acting in accordance with it, viewing it as unproblematic, etc.
Since there can be many criticisms (which are also just ideas) and deeply nested counter-criticisms, the result is a tree structure. For example, as a discussion progresses, one of its trees might look like this:
I/ | \C11 C12 C13/ \ \C21 C22 C23/ \C31 C32
In this tree, I is considered problematic. Although C11 has been neutralized by C21 and C22, C12 still needs to be addressed. In addition, C23 would have neutralized C13, but C31 and C32 make C23 problematic, so C13 makes I problematic as well.
You don’t need to keep track of these relationships manually. Veritula automatically marks ideas accordingly.
Since decision-making follows the same logic as truth-seeking, you can use these trees to make decisions, too. Veritula implements unanimous consent as defined by Taking Children Seriously, a parenting philosophy that builds on Popper’s epistemology. When you’re planning your next move but can’t decide on a city, say, Veritula helps you criticize your ideas and make a rational decision – meaning a decision you’ll be happy with. Again, it’s rational to act in accordance with ideas that have no pending criticisms.
All ideas, including criticisms, should be formulated as concisely as possible, and separate ideas should be submitted separately, even if they’re related. Otherwise, you run the risk of receiving ‘bulk’ criticisms, where a single criticism seems to apply to more content than it actually does.
Again, criticisms are also just ideas, so the same is true for criticisms. Submitting each criticism separately has the benefit of requiring the proponent of an idea to address each criticism individually, not in bulk. If he fails to address even a single criticism, the idea remains problematic and should be rejected.
The more you discuss a given topic, the deeper and wider the tree grows. Some criticisms can apply to multiple ideas in the tree, but that needs to be made explicit by submitting them repeatedly.
Comments that aren’t criticisms – eg follow-up questions or otherwise neutral comments – are considered ancillary ideas. Unlike criticisms, ancillary ideas do not invert their respective parents’ statuses. They are neutral.
One of the main benefits of Veritula is that the status of any idea in a discussion can be seen at a glance. If you are new to a much-discussed topic, adopt the displayed status of the ideas involved: if they are marked problematic, reject them; if they are not, adopt them.
Therefore, Veritula acts as a dictionary for ideas.
One of the problems of our age is that people have same discussions over and over again. Part of the reason is widespread irrationality, expressed in the unwillingness to change one’s mind; another is that it’s simply difficult to remember or know what’s true and what isn’t. Discussion trees can get complex, so people shouldn’t blindly trust their judgment of whether some idea is true or problematic, whether nested criticisms have been neutralized or not. Going off of memory is too error prone.
Veritula solves this problem: it makes discussion trees explicit so you don’t have to remember each idea and its relation to other ideas. Veritula therefore also enables you to hold irrational people accountable: if an idea has pending criticisms, the rational approach is to either abandon it or to save it by revising it or addressing all pending criticisms.
Many people don’t like to concede an argument. But with Veritula, no concessions are necessary. The site just shows you who’s right.
Using Veritula, we may discover a bit of truth.
Still have questions about rationality? Read the technical paper, learn about Veritula’s recursive epistemology and how to structure discussions, or read the quick guide to rational decision-making.
Want personalized help tailored to your own specific needs? Hire Dennis, founder of Veritula, for personal tutoring, or run a bounty.
Popperian epistemology has some flaws, like verisimilitude, but Veritula doesn’t implement those.
A simple implementation of hard-to-vary
This is my submission for Veritula's bounty for Idea #3069, which asks for an executable implementation that can compare arbitrary English explanations by how hard they are to vary.
Code: hard-to-vary
I've implemented the approach as a small interactive program. It takes arbitrary explanations supplied by the user, collects working variations of each, and ranks the explanations by hardness to vary.
I think hard-to-vary can be implemented more simply than the approaches considered in Dennis's blog, Hard to Vary or Hardly Usable?. The basic idea I'm using is this: if an explanation is harder to vary, there should be fewer ways of changing it while still having it explain what it's supposed to explain.
The program starts by asking the user what question they're trying to answer and then asks for a list of proposed explanations. It goes through each explanation one at a time and asks the user to come up with variations of it that would still work as explanations. The user can enter as many as they can find, then move on to the next explanation.
Finally, the program counts the working variations and ranks the explanations. Fewer working variations means harder to vary, while equal numbers mean equally hard to vary. So if explanation A has two working variations and explanation B has five, A is harder to vary than B.
An example
Take the question: Why do seasons occur?
One explanation is the story of Persephone. She spends part of the year in Hades, Demeter grieves while she's gone, and this accounts for the seasonal cycle.
I might try changing the explanation in a few ways:
- Persephone spends five months in Hades instead of six.
- It's Demeter's son who is taken rather than her daughter.
- Zeus orders Persephone to return periodically rather than her return depending on the details involving the pomegranate.
The details have changed, but the basic explanatory story can still do the same job. Suppose I find three such variations.
Now I try the explanation involving Earth's axial tilt. I might be able to change the stated tilt from 23.4 degrees to 23.5 degrees without substantially affecting the explanation. But if I start making large changes to the geometry while trying to explain the same observations, the explanation stops working. Suppose I find one working variation.
The program therefore ranks axial tilt as harder to vary: one working variation versus three. That's the entire comparison procedure.
How I understand Dennis's criticism
I don't understand Dennis as claiming that Deutsch simply gets the Persephone example wrong. The issue in the blog is how we get from examples like this, where we seem to have an intuition that one explanation is harder to vary, to a sufficiently specified procedure that could compare explanations in general.
Dennis initially explores numerical quality scores, but that immediately creates problems. Why should one explanation have a score of 500 rather than 550? Why choose that scale? How do criticisms affect the score? How do criticisms of criticisms affect it? I agree with Dennis that these choices look arbitrary.
He eventually gets rid of the quality scores entirely. Instead of trying to measure the quality of an idea, his program keeps track of pending criticisms. His proposed rule becomes: adopt ideas without pending criticisms and reject ideas that have them.
There's something important about how that system works, though. The program doesn't generate criticisms itself; people do. And that's intentional: Dennis says "I’m not looking to formalize or automate creativity as a whole". Creative input can come from users while the program handles the non-creative part of the process. I agree that you can have a rational decision-making process while outsourcing the creative part to the user.
I just don't see why we can't do the same thing with HTV. Let the user come up with variations and say which ones they think still work. The program doesn't need to understand the explanation or come up with the variations itself. It keeps track of the variations and compares the counts.
This doesn't seem fundamentally different from Dennis letting the user tell his program that something is a criticism. In fact, in the blog he explicitly avoids having the program figure out whether a comment is really a criticism: the user checks a box saying that it is. So in my program the user is supplying a working variation; in his, the user is supplying a criticism. In both cases the user is providing the part that requires understanding and judgment, and the program does something simple with that input.
It also means I don't need the quality scores Dennis runs into trouble with. I don't need to decide how many points axial tilt gets compared with Persephone. I'm just asking: how many ways have we actually found to change each explanation while still having it work?
But isn't this subjective?
One criticism Dennis quotes is:
“Also, isn’t the difficulty of changing an explanation at least partly a property not of the explanation itself but of whoever is trying to change it? If I’m having difficulty changing it, maybe that’s because I lack imagination. Or maybe I’m just new to that field and an expert could easily change it.”
I think this is true. I just don't think it's a problem specific to HTV.
Imagine I can't think of any working variations of an explanation, but an expert can immediately think of five. Then yes, our results will be different. But isn't that also what happens with criticism? I might look at an idea and fail to see anything wrong with it while someone who knows much more about the subject immediately sees a serious criticism.
The same goes for participation. An idea on Veritula might have zero pending criticisms simply because hardly anyone has tried to criticize it. That doesn't mean there are literally no criticisms of it. Someone could find one tomorrow. Dennis's answer in the blog is basically that if this bothers you, try to find a criticism yourself. If you can't find one, why not adopt the idea?
I think HTV can work the same way. Zero working variations doesn't mean that we've somehow proven there are no possible variations. It means we haven't found one. If you think the explanation is actually easy to vary, try to come up with a variation that still works.
So yes, the result depends on the knowledge and creativity of the person using the program. But I think rational decision-making is always going to depend on what criticisms, arguments, alternatives, etc. a person is actually aware of. I don't see how Veritula escapes that either.
What about human judgment?
There's another obvious question: who decides whether a variation actually works?
For this program, the user does. I don't think we can get rid of that kind of human judgment, at least until we get AGI. Two people can disagree about whether a variation still explains the thing we're trying to explain. They can also disagree about whether two variations are really different or are basically the same variation stated twice.
Again, I think Veritula has the same underlying issue. People still have to decide whether something really is a criticism, whether a countercriticism actually answers it, whether two criticisms are redundant, and so on. Dennis's program can keep track of the structure, but the structure only means something if those judgments make sense.
Take an extreme case. If a malicious moderator rejects every good criticism of an idea and accepts nonsense countercriticisms, the idea could end up showing 0 pending criticisms. I obviously shouldn't look at the 0 and conclude that the idea is rational to adopt. I'd want to read what happened and decide whether I agree with it.
I don't mean this as a criticism specific to Veritula. I think it's just a limit of this kind of approach. At some point people have to make judgments, and people can disagree about them. Ultimately everyone is their own moderator when it comes to their own decision-making. I have to decide which arguments I accept, which criticisms I think have been answered, and so on.
The same is true with my HTV program. If someone gives me a ranking based on ten supposed working variations, I don't have to accept the ranking blindly. I can look at the ten variations and decide that five don't really work and three others are basically duplicates. My result would then be different.
I'm fine with that. I don't think the goal of either program should be to somehow remove judgment from rational thinking. The program gives us a procedure for what to do with the judgments we've made.
What does the program actually contribute?
Dennis writes:
“We can’t just outsource everything to the user – the app has to do some things or it has no value.”
This was actually the part of the blog that made me think about Veritula itself. What is Veritula doing, and what is it outsourcing?
It outsources the interesting creative part to people. People come up with the ideas. People come up with the criticisms. The user can even tell the program whether something they've written is a criticism by checking a box. The program then keeps track of the structure and tells us how many criticisms are pending.
My program is doing something similar, except with variations. People come up with the explanations and the working variations. The program keeps track of them, counts them, and ranks the explanations.
So I don't think I'm outsourcing everything to the user any more than Veritula is. I'm outsourcing the part that requires creativity and judgment. The actual decision rule is implemented in the program.
For Veritula, that rule ultimately depends on whether there are pending criticisms. For my program, it depends on the number of working variations: fewer working variations means harder to vary.
Where I disagree with the blog
I agree with Dennis that the quality sliders in the blog don't work. I also agree with his decision to let users supply the creative input rather than expecting the program to generate it.
Where I disagree is that I think once we allow this same freedom for HTV, we can construct a similarly simple program for it. The user comes up with explanations and tries to vary them while keeping them working. The program counts the working variations and ranks the explanations.
There are still all the normal problems of human knowledge: maybe I missed a variation, maybe I accepted a bad one, maybe someone else would judge things differently. But those same problems exist when we come up with and judge criticisms.
I don't think either program solves those problems, and I don't think it needs to. That's the part people do.
There’s currently no validation, neither automatic nor creative. A user could input junk and there’s no way to correct that short of running the program a second time. (The revision may already address this criticism, I haven’t checked it yet.)
… nor is the user given an opportunity to correct any mistakes to that effect.
The user could simply run the program again with corrected inputs.
There are still all the normal problems of human knowledge: maybe I missed a variation, maybe I accepted a bad one, maybe someone else would judge things differently. But those same problems exist when we come up with and judge criticisms.
In this context, there’s a serious issue at the heart of HTV in that not all preference formation involves thinking about variants. Oftentimes, preference formation simply doesn’t take this form. Successful rehabilitation of HTV would be universal in this sense: it would reduce all preference formation to HTV somehow.
I’ve given the example of Pop Tarts in the past. You can ask someone why they’re called that. And they might guess some answers (‘because they “pop” out of the toaster’, ‘because the name is fun’, ‘because the filling tastes tart’, etc.). But ultimately the preferable (and correct) answer will be gotten by just looking it up on the company website, say. Not by seeing which guess is “harder to vary”. I think that should count toward the “normal problems” you mention, but it’s one that needs to be addressed.
➜ hard-to-vary git:(master) python3 htv.pyHard-to-Vary Explanation Comparator-----------------------------------What question are you trying to answer?> Why are they called Pop Tarts?Enter explanations one at a time.Press Enter on an empty line when you're done.Explanation 1: Because they ‘pop’ out of the toaster.Explanation 2: Because the name is fun.Explanation 3: Because the filling tastes tart.Explanation 4:============================================================Explanation:Because they ‘pop’ out of the toaster.Enter variations of this explanation that would still work to explain the same thing.Press Enter on an empty line when you're done.Variation 1:============================================================Explanation:Because the name is fun.Enter variations of this explanation that would still work to explain the same thing.Press Enter on an empty line when you're done.Variation 1:============================================================Explanation:Because the filling tastes tart.Enter variations of this explanation that would still work to explain the same thing.Press Enter on an empty line when you're done.Variation 1:============================================================Question:Why are they called Pop Tarts?============================================================HARDNESS-TO-VARY RANKING============================================================Rank 1Explanation: Because they ‘pop’ out of the toaster.Working variations submitted: 0 variationsRank 1Explanation: Because the name is fun.Working variations submitted: 0 variationsRank 1Explanation: Because the filling tastes tart.Working variations submitted: 0 variationsFewer working variations = harder to vary.
In this case, I couldn’t think of any good-faith variants of any of the three guesses. So they end up being equally preferable according to the program. But I happen to know that all three are false.
If I didn’t know, I’d reject guess 3 because Pop Tarts don’t actually taste tart. They taste sweet. So that’s a pending criticism. But guess 1 and 2 are plausible and I see no reason to prefer one over the other. One would need additional info, but not variant guesses.
I don't think either program solves those problems, and I don't think it needs to. That's the part people do.
I think Veritula solves it because it has a recursive notion of criticism. It doesn’t ‘solve’ it in the sense that it completely takes all burden off humans – creative input is still required – but it solves it in the sense that new input can be given at runtime that changes the displayed (tentative) result.
In any case, although the blog post suggests Veritula as a working alternative to HTV, the bounty isn’t a competition between HTV and Veritula. It’s about rehabilitating HTV in its own right. That may require addressing flaws even if Veritula has them, too.
There are still all the normal problems of human knowledge: maybe I missed a variation, maybe I accepted a bad one, maybe someone else would judge things differently. But those same problems exist when we come up with and judge criticisms.
That’s true, but it isn’t clear how criticisms fit into the program as proposed. The BoI chapter 1 glossary (p. 31) defines:
Good/bad explanation An explanation that is hard/easy to vary while
still accounting for what it purports to account for.
The program currently has no way to tell, neither automatically nor through human input, whether a variation of an explanation still accounts for what it purports to account for. One could run the program again with corrected inputs but as I noted in #5533, that approach won’t scale.
When variations are given for both competing explanations, the user is none the wiser which variation of the ‘better’ explanation he should adopt. There’s no clear preference:
Whenever a wide range of variant theories can account equally well for the phenomenon they are trying to explain, there is no reason to prefer one of them over the others, so advocating a particular one in preference to the others is irrational.
#5526 is an example. Even ignoring the obviously false outcome, which I caused for illustrative purposes, we don’t know which variation of the better explanation to prefer.
… nor is the user given an opportunity to correct any mistakes to that effect.
The user could simply run the program again.
@dirk-meulenbelt says it isn’t clear what counts as a variation vs an entirely different explanation. I’ll add that the program can’t currently recognize the difference, nor is the user given an opportunity to correct any mistakes to that effect.