What we can and cannot say about these seven questions
What we can and cannot say about these seven questions
What we can and cannot say about these seven questions
Seven real questions from Whetstone’s opening week, the design decision behind each one, and an honest account of where the evidence runs out.
Most products in this space make a promise they cannot support. Play these games, get smarter.
In January 2016 the Federal Trade Commission settled deceptive advertising charges against Lumos Labs, the maker of Lumosity. The company paid $2 million. A $50 million judgment was entered and the remainder was suspended because of the company’s financial condition. Jessica Rich, then Director of the FTC’s Bureau of Consumer Protection, said the company “preyed on consumers’ fears about age-related cognitive decline” and “simply did not have the science to back up its ads.” For claims that a product helps performance at school, at work or in athletics, or holds off age-related decline, that order requires human clinical testing that is randomized, adequately controlled, and blinded to the maximum extent practicable. We hold nothing of the kind for anything on this page.
In October 2014 the Stanford Center on Longevity and the Berlin Max Planck Institute for Human Development published a consensus statement signed by more than 70 cognitive psychologists and neuroscientists, stating that the literature does not support claims that software games alter neural functioning in ways that carry into everyday performance, or prevent decline. That statement is about commercial software. Its silence about anything else is silence, not approval.
I am telling you this before I tell you anything about what I built, because that is the standard I have to hold myself to, and almost nobody in this category does.
So, plainly, before you read another line: we have no evidence that any of this does anything to your thinking, we are not going to run the trials that could tell us, and nothing below should be read as a claim that it does.
What follows is the opening week of Whetstone, one question at a time, with what each one is built to do and where the evidence stops. If you finish it and decide you do not need the product, that is a fine outcome. The seven questions are printed in full and you can do all of them yourself, today, for nothing.
Seven real questions from Whetstone’s opening week, the design decision behind each one, and an honest account of where the evidence runs out.
Most products in this space make a promise they cannot support. Play these games, get smarter.
In January 2016 the Federal Trade Commission settled deceptive advertising charges against Lumos Labs, the maker of Lumosity. The company paid $2 million. A $50 million judgment was entered and the remainder was suspended because of the company’s financial condition. Jessica Rich, then Director of the FTC’s Bureau of Consumer Protection, said the company “preyed on consumers’ fears about age-related cognitive decline” and “simply did not have the science to back up its ads.” For claims that a product helps performance at school, at work or in athletics, or holds off age-related decline, that order requires human clinical testing that is randomized, adequately controlled, and blinded to the maximum extent practicable. We hold nothing of the kind for anything on this page.
In October 2014 the Stanford Center on Longevity and the Berlin Max Planck Institute for Human Development published a consensus statement signed by more than 70 cognitive psychologists and neuroscientists, stating that the literature does not support claims that software games alter neural functioning in ways that carry into everyday performance, or prevent decline. That statement is about commercial software. Its silence about anything else is silence, not approval.
I am telling you this before I tell you anything about what I built, because that is the standard I have to hold myself to, and almost nobody in this category does.
So, plainly, before you read another line: we have no evidence that any of this does anything to your thinking, we are not going to run the trials that could tell us, and nothing below should be read as a claim that it does.
What follows is the opening week of Whetstone, one question at a time, with what each one is built to do and where the evidence stops. If you finish it and decide you do not need the product, that is a fine outcome. The seven questions are printed in full and you can do all of them yourself, today, for nothing.
One study worth knowing about
In a 2025 CHI paper, researchers at Microsoft Research and Carnegie Mellon surveyed 319 knowledge workers about 936 real tasks where they used generative AI. Higher confidence in the AI was associated with less self-reported critical thinking. Higher confidence in their own ability went the other way.
The tempting read is “AI is making people worse at thinking.” That is not what the paper shows, and I am not going to say it. These are self-reported perceptions from a one-time survey. They show association, not cause, and nothing in that paper says any practice changes what it measured.
The accurate version is smaller and more interesting: the more you trust the model, the less you check it. The more you trust yourself, the more you do.
A note on what is actually built here
Four things happen in this product. I can only defend one of them with research, so here is the split before you read the questions.
It removes the decision of what to reflect on. No research. That is a design argument.
It hands you a question you would not have chosen. No research. Also a design argument.
It makes you finish the thought in writing. This one has support. Material you generate yourself is remembered better than material you only read. That finding is well replicated. It is also narrow: it is about memory for the thing you generated, not about thinking in general, and I am not going to stretch it past that.
It collects what you wrote into a printed object at the end of the month. No research at all.
Three of the four are design arguments. I would rather say that here than imply otherwise for seven questions in a row.
1. Where They Actually Stand
Clear a patch of table or floor. Put one object down for yourself, then one for each of the five or six people who take up the most room in your life right now. Coins, keys, cups, whatever is within reach. Place each one at the distance that is true, from you and from each other, not the distance you would describe out loud. Someone can be important and far. Someone can be close and not chosen. Sit back and read the shape you just made, especially the part you did not expect to see. Then move one object to where you want it by the end of the year, and name the first move that is actually yours to make.
The move. Put an invisible structure outside your head, where you have to look at it.
Why it is built this way. Day one has one job, which is to get answered, so it uses objects already within reach and no tools. The design decision is the phrase “the distance that is true, not the distance you would describe out loud.” Asked in words, you produce the socially correct answer, because words let you hedge. A coin is somewhere or it is not. The second decision is the last line. Reading the shape is interesting and changes nothing, so the question ends by demanding one move that does not depend on the other person doing anything first.
What it does not show. We have no evidence that arranging objects tells you anything true about your relationships, and none that doing it changes them. You may lay it out and feel nothing. What it reliably does is force one specific commitment out of a vague feeling, on a day when a vague feeling was the alternative.
2. The Twenty Seconds You’d Take
Think of the last time you felt envy, even a flicker at someone’s news. Do not take their whole life. Name the exact twenty seconds of it you would take. Now split that moment three ways: is it the thing itself, the ease with which they have it, or what having it would prove about you? Then the test: would you still want it if nobody ever knew you had it? Now run the whole thing again on a second flicker. Which of the three wins both times?
The move. Take one feeling apart into named components, then check which component survives a test.
Why it is built this way. Envy is usually discussed as something to be talked out of. This question does not do that, because being talked out of a feeling teaches you nothing about it. The twenty second limit is doing the real work: “I envy their life” is unanswerable, and “I envy the twenty seconds where they told everyone” is answerable. The secrecy test is there because it separates wanting a thing from wanting to be seen having it, and those two feel identical from the inside. The second run is not padding. One data point is a mood, and two is the beginning of a pattern.
What it does not show. No claim that this reduces envy, changes what you want, or makes you happier. We have not tested it and I do not know of anyone who has. It is one deliberate look at a feeling most people only glance at.
3. The Wordless Manual
Pick one thing you do well that nobody taught you: how you start work you are dreading, how you come down after an argument, how you get yourself to sleep when your head is loud. Draw it as the kind of manual that comes with furniture you assemble yourself. Numbered panels, arrows, hands doing things, no words anywhere. Then find the panel a stranger would still get wrong, and split it into the steps you never had to say out loud.
The move. Convert something you do automatically into instructions somebody else could follow.
Why it is built this way. The no words rule is the entire design. Given words, you will write “calm down” and believe you have described something. Given only panels and arrows, you cannot hide behind an abstraction, because you have to draw what the hands actually do. The last instruction is the point of the exercise: the panel a stranger would get wrong is the step you know so well you stopped seeing it, and that step is usually the one that makes the whole thing work. Drawing badly is fine and is not scored, because the difficulty is meant to be in the specifying, not in the draughtsmanship.
What it does not show. No claim that this improves how you explain things generally, or that the thing you drew will work for anyone else. We have not tested that. It is one attempt to say precisely something you have only ever done.
4. The Times Left
Pick one person you love who you do not live with. Guess in one second: how many more times will you two be in the same room, ever? Write it down. Now build the real number. Count the actual times you were together last year, month by month, not a vibe. That is your rate. Multiply it by the years you both realistically have, the shorter of the two. Compare that total to your guess. Then change the one number you control: add one occasion a year, every year, and rerun it. What does one more buy you?
The move. Replace an impression with an arithmetic estimate, then test what one variable changes.
Why it is built this way. The one second guess is captured before the arithmetic on purpose, because once you have built the real number you can no longer remember what you assumed, and the gap between the two is the entire exercise. Counting month by month rather than estimating a rate is deliberate, since a remembered rate is a feeling and a counted one is not. The final step exists because a number like this can leave you with nothing but dread, and dread is not useful. Adding one occasion a year is the only variable in the calculation you actually control.
What it does not show. Nothing here says you will act on the number, and nothing says people who do this see anyone more often. This is also not a prediction about anyone’s life expectancy, and it should not be treated as one. It is an estimate built from your own inputs, and its only claim on you is that you built it.
5. The Good Clothes
Go and get the thing you never use because you are saving it: the jacket bought for a life you were expecting, the good shoes, the pen, the notebook too nice to write in. Use it right now on the most ordinary thing it could possibly do, the jacket on while you wash a plate, tomorrow’s list in the notebook. Then say out loud what you were saving it for, the actual occasion you had in mind, and whether that occasion still exists in your life. Rule on it: a real date in the diary, ordinary use starting this week, or out.
The move. Act first, then examine the belief the action exposed.
Why it is built this way. The order is the design. If you ask the question first, you get a reasonable sounding answer about wanting to look after nice things. If you make somebody wear the good jacket to wash a plate first, the question lands on a person who has just felt something, usually a small resistance they were not expecting. Saying the occasion out loud matters, because most of these occasions turn out to belong to a version of your life that either already happened or is not coming. The three way ruling exists so the question cannot be enjoyed and then dropped, and one of the three options is deliberately to keep saving it, as long as you name a date.
What it does not show. No claim about your relationship with money or objects, and nothing here is a diagnosis of anything. We think most people are saving things for a life they are no longer living. We have not tested that, and neither has anyone else that we can find.
6. The Dish That Outlives You
Invent the dish or the drink that your family, or whoever you count as yours, keeps for a hundred years. It does not exist yet. Decide what it is, the occasion it belongs to, and the one rule about how it gets served that nobody is allowed to break. Now write it down the way it would have to be written to survive without you, for someone who has never seen it and cannot ask you a single question. Read your own instructions back cold, find the step where that person gets it wrong, and fix it.
The move. Write instructions for a reader who cannot ask you anything, then find where they break.
Why it is built this way. The hundred year frame is not sentiment, it is a specification: it removes you from the room. Every set of instructions any of us writes quietly assumes a reader who can ask a follow up question, and this one forbids that. The unbreakable rule is there because traditions survive on their constraint rather than on their content, and picking the constraint forces you to say what the thing is actually for. The final step, reading it back cold and finding where a stranger goes wrong, is the same operation as rep 3 pointed at something you invented rather than something you already do.
What it does not show. No claim that this makes you a better writer, that your family will adopt anything, or that traditions can be manufactured. We do not know whether they can. The exercise is the specifying, and the dish is the excuse for it.
7. The Wear Marks
Before you get up, write down the three objects in the place you live that you think carry the most wear from your own hands. Now walk it and find the real ones: the shine, the dent, the frayed patch, the worn key, the cracked corner. Not the oldest things, the most used. Then find the one thing sitting almost untouched that you bought or kept because you were going to use it. Put your three guesses next to what you actually found. Then write the two sentences a stranger could write using only the wear as evidence: one about where your hours really go, one about who you were planning to be.
The move. Predict, then check against physical evidence you left without meaning to.
Why it is built this way. Guessing before walking is the whole design, because without the guess this is a pleasant tour of your home and with it there is something at stake. Wear is a better witness than memory: it cannot be edited afterwards and it accumulated while you were not paying attention. The untouched object is the counterpart of the worn ones, and it is doing separate work, because the gap between what you use and what you bought to use is where intention and behaviour come apart. Closing the week here is deliberate, since day one asked you to place people by a truth you would not say out loud, and this asks the objects to say it instead.
What research supports. Strangers who saw only a person’s bedroom or office, and never met them, formed impressions that partly matched how that person and their friends described them. The match was clearest for openness and conscientiousness, and near zero for agreeableness in offices. In the same studies, observers also read meaning into room features that did not track the occupant at all, and part of the agreement between strangers ran through assumptions about sex and race. A room is a weak signal, not a readout. *(Gosling, Ko, Mannarelli & Morris, 2002, Journal of Personality and Social Psychology. Tier B: supported, narrow.)*
What it does not show. That work measured strangers reading rooms. It did not measure whether you can read your own, which is the thing this question asks. We are not claiming improved self-knowledge or improved social perception. We are claiming one deliberate look at evidence you have been leaving around without deciding to.
The closing
Here is what a week of this does not do. It does not raise your IQ. It does not make you measurably better at reasoning in general. Almost nothing does, the evidence for far transfer is weak across the board, and the largest company that said otherwise paid $2 million to stop saying it. If someone sells you a daily practice on the promise of general cognitive improvement, they are selling you something that has not been demonstrated and that they would struggle to defend if anyone asked them to.
Here is what it does do, and this part we are confident about because it is a description rather than a prediction. Seven days from now you will have written seven things you would not otherwise have written. Each one came from a question you did not choose, which is the only reason they went anywhere unfamiliar.
That is the whole offer. One question a day that you did not pick, answered in your own words, and at the end of the month what you wrote comes back to you printed on paper.
If that sounds worth fifteen dollars, it is at usethriv.com/whetstone. If it sounds like something you would rather do yourself with a notes app and an alarm, do that. The questions are above.
Where every claim above comes from
Every source, and what each one does not reach. Two of the seven questions cite research; five cite nothing, because nothing studies them.
Opening. FTC settled with Lumos Labs, $2m paid, $50m entered and suspended; RCT standard for school/work/athletics and age-related decline claims
Source. FTC press release and order, 5 Jan 2016
What it does not reach. The order binds Lumos Labs, not us. A separate, lower standard covers other efficacy claims
Opening. More than 70 scientists, software games do not carry into everyday performance
Source. Stanford Center on Longevity / Max Planck consensus, Oct 2014
What it does not reach. It is about commercial software. It says nothing for or against anything else
Opening. Training on a task improves that task, with little evidence of carryover
Source. Simons et al. (2016), *Psychological Science in the Public Interest* 17(3), 103-186
What it does not reach. The same standard applies to us, and we do not meet it
One study. 319 knowledge workers, 936 tasks, confidence in AI vs confidence in self
Source. Lee, Tankelevitch et al. (2025), CHI
What it does not reach. Self-reported, correlational, one-time survey. No practice was tested
What is built. Material you generate is remembered better than material you read
Source. Generation effect, Slamecka & Graf (1978) and subsequent meta-analysis
What it does not reach. Memory for the generated material only. Says nothing about thinking in general
Rep 1. None
Source. None
What it does not reach. Nothing studies this. Design argument only
Rep 2. None
Source. None
What it does not reach. Nothing studies this. Design argument only
Rep 3. None
Source. None
What it does not reach. Nothing studies this. Design argument only
Rep 4. None
Source. None
What it does not reach. Nothing studies this. Design argument only
Rep 5. None
Source. None
What it does not reach. Tier D admission printed in the rep itself
Rep 6. None
Source. None
What it does not reach. Nothing studies this. Design argument only
Rep 7. Strangers reading a room form partly accurate impressions, and also read meaning into features that carry none
Source. Gosling, Ko, Mannarelli & Morris (2002), *JPSP*
What it does not reach. Does not measure whether you can read your own room