Whetstone

What we can and cannot say about these seven questions

What we can and cannot say about these seven questions

What we can and cannot say about these seven questions

Seven real questions from Whetstone’s opening week, the design decision behind each one, and an honest account of where the evidence runs out.
Most products in this space make a promise they cannot support. Play these games, get smarter.
In January 2016 the Federal Trade Commission settled deceptive advertising charges against Lumos Labs, the maker of Lumosity. The company paid $2 million. A $50 million judgment was entered and the remainder was suspended because of the company’s financial condition. Jessica Rich, then Director of the FTC’s Bureau of Consumer Protection, said the company “preyed on consumers’ fears about age-related cognitive decline” and “simply did not have the science to back up its ads.” For claims that a product helps performance at school, at work or in athletics, or holds off age-related decline, that order requires human clinical testing that is randomized, adequately controlled, and blinded to the maximum extent practicable. We hold nothing of the kind for anything on this page.
In October 2014 the Stanford Center on Longevity and the Berlin Max Planck Institute for Human Development published a consensus statement signed by more than 70 cognitive psychologists and neuroscientists, stating that the literature does not support claims that software games alter neural functioning in ways that carry into everyday performance, or prevent decline. That statement is about commercial software. Its silence about anything else is silence, not approval.
I am telling you this before I tell you anything about what I built, because that is the standard I have to hold myself to, and almost nobody in this category does.
So, plainly, before you read another line: we have no evidence that any of this does anything to your thinking, we are not going to run the trials that could tell us, and nothing below should be read as a claim that it does.
What follows is the opening week of Whetstone, one question at a time, with what each one is built to do and where the evidence stops. If you finish it and decide you do not need the product, that is a fine outcome. The seven questions are printed in full and you can do all of them yourself, today, for nothing.
Seven real questions from Whetstone’s opening week, the design decision behind each one, and an honest account of where the evidence runs out.
Most products in this space make a promise they cannot support. Play these games, get smarter.
In January 2016 the Federal Trade Commission settled deceptive advertising charges against Lumos Labs, the maker of Lumosity. The company paid $2 million. A $50 million judgment was entered and the remainder was suspended because of the company’s financial condition. Jessica Rich, then Director of the FTC’s Bureau of Consumer Protection, said the company “preyed on consumers’ fears about age-related cognitive decline” and “simply did not have the science to back up its ads.” For claims that a product helps performance at school, at work or in athletics, or holds off age-related decline, that order requires human clinical testing that is randomized, adequately controlled, and blinded to the maximum extent practicable. We hold nothing of the kind for anything on this page.
In October 2014 the Stanford Center on Longevity and the Berlin Max Planck Institute for Human Development published a consensus statement signed by more than 70 cognitive psychologists and neuroscientists, stating that the literature does not support claims that software games alter neural functioning in ways that carry into everyday performance, or prevent decline. That statement is about commercial software. Its silence about anything else is silence, not approval.
I am telling you this before I tell you anything about what I built, because that is the standard I have to hold myself to, and almost nobody in this category does.
So, plainly, before you read another line: we have no evidence that any of this does anything to your thinking, we are not going to run the trials that could tell us, and nothing below should be read as a claim that it does.
What follows is the opening week of Whetstone, one question at a time, with what each one is built to do and where the evidence stops. If you finish it and decide you do not need the product, that is a fine outcome. The seven questions are printed in full and you can do all of them yourself, today, for nothing.

One study worth knowing about

In a 2025 CHI paper, researchers at Microsoft Research and Carnegie Mellon surveyed 319 knowledge workers about 936 real tasks where they used generative AI. Higher confidence in the AI was associated with less self-reported critical thinking. Higher confidence in their own ability went the other way.
The tempting read is “AI is making people worse at thinking.” That is not what the paper shows, and I am not going to say it. These are self-reported perceptions from a one-time survey. They show association, not cause, and nothing in that paper says any practice changes what it measured.
The accurate version is smaller and more interesting: the more you trust the model, the less you check it. The more you trust yourself, the more you do.

A note on what is actually built here

Four things happen in this product. I can only defend one of them with research, so here is the split before you read the questions.
It removes the decision of what to reflect on. No research. That is a design argument.
It hands you a question you would not have chosen. No research. Also a design argument.
It makes you finish the thought in writing. This one has support. Material you generate yourself is remembered better than material you only read. That finding is well replicated. It is also narrow: it is about memory for the thing you generated, not about thinking in general, and I am not going to stretch it past that.
It collects what you wrote into a printed object at the end of the month. No research at all.
Three of the four are design arguments. I would rather say that here than imply otherwise for seven questions in a row.

1. Where They Actually Stand

Write down the five or six people who take up the most room in your life right now. Rank them once by closeness, your honest order, before you touch your phone. Then predict a second order: how recently you last had a real exchange with each, message or call, either direction. Now check the phone, one date per person, last real contact only, no counting. Write the true order beside your guess. Wherever the two orders disagree, sit with the gap: a person can slide through drift, or because somebody decided something. For each gap, say which it is, and if it was a decision, whose.
The move. Rank what you believe, then check it against the record your phone kept without asking you.
Why it is built this way. Day one has one job, which is to get answered, so everything in it is bounded: five or six people, one date each, no counting. The ranking goes on paper before the phone comes out, and the order matters because the phone ends the argument. The dates overwrite whatever you would have guessed, and a guess that was never written down quietly becomes a guess that was always right. The closeness ranking is what turns the dates into information: a slow month with an acquaintance is nothing, and a slow month with the person you put second is a finding. “A real exchange, either direction” is a definition with a job: it stops a like from counting as contact, and it stops you from counting only what other people owe you. The ending is there because a gap can be read and shrugged at. Drift or a decision, and whose, is the difference between noticing something and being responsible for it.
What it does not show. We have no evidence that how recently you spoke to someone measures what they mean to you, and the exercise does not assume it does. It asks about the disagreement between what you ranked and what the dates say, and the explanation is yours to give, not ours to score. No claim that seeing it changes anything about your relationships. What you get is one impression about your own life checked against a record you did not keep on purpose, on a day when the impression was the alternative.

2. The Twenty Seconds You’d Take

Think of the last time you felt envy, even a flicker at someone’s news. Do not take their whole life. Name the exact twenty seconds of it you would take. Now split that moment three ways: is it the thing itself, the ease with which they have it, or what having it would prove about you? Search your own last year for the nearest thing to those twenty seconds, a moment somebody could have envied you for. Find a real one. Were you in it when it happened, or somewhere else in your head? Now look at their twenty seconds again and say which of the three you are actually short of.
The move. Take one feeling apart into named components, then aim the same split at yourself.
Why it is built this way. Envy is usually discussed as something to be talked out of. This question does not do that, because being talked out of a feeling teaches you nothing about it. The twenty second limit is doing the real work: “I envy their life” is unanswerable, and “I envy the twenty seconds where they told everyone” is answerable. The mirror is the second test, and it is not there to comfort you. It exists for the question that follows it: were you in it when it happened? If you were somewhere else in your head, you have learned something about the inside of enviable moments that looking at theirs cannot teach you. The last line turns the feeling into an inventory. Thing, ease, or proof: you are usually short of one, and it is rarely the thing.
What it does not show. No claim that this reduces envy, changes what you want, or makes you happier. There is no study behind this question, ours or anyone’s, and writing carefully around that would just be a quieter way of lying. Your one remembered moment is not evidence about anyone else’s, and we are not pretending it is. It is one deliberate look at a feeling most people only glance at.

3. The Steps You Never Said

Pick one thing you do well that nobody taught you: how you start work you are dreading, how you come down after an argument, how you get yourself to sleep when your head is loud. Write it as a numbered procedure, each step an instruction you could hand to someone, no step wider than one action. Now test it against the record. Recall the last time it worked and walk your steps against what you actually did that day. Then recall a time it failed. Somewhere your procedure and the failure disagree: a step you skipped, or one you do without knowing it is a step. Add it. That is the skill.
The move. Convert something you do automatically into steps, then audit the steps against what actually happened.
Why it is built this way. The one-action rule is where the honesty comes from. Given a wider step, you will write “calm down” and believe you have described something. Held to one action per step, you cannot hide behind an abstraction, because you have to write what the hands and the minutes actually do. Then the procedure meets the record, and the order of the two memories matters: the day it worked calibrates the steps, and the day it failed interrogates them. The disagreement has only two shapes, a step you skipped or a step you never knew you were taking, and both measure the same thing: the distance between the skill and your description of it. Closing that distance is what the prompt’s last sentence names.
What it does not show. Whether this improves how you explain anything else, we do not know and have not tested. A procedure reconciled against two remembered days is not a study of your own skill, and the added step does not become true by being written down. It leaves you with a procedure one failure more complete than it was, and one attempt to say precisely something you have only ever done.

4. The Times Left

Pick one person you love who you do not live with. Guess in one second: how many more times will you two be in the same room, ever? Write it down. Now build the real number. Count the actual times you were together last year, month by month, not a vibe. That is your rate. Multiply it by the years you both realistically have, the shorter of the two. Compare that total to your guess. Then change the one number you control: add one occasion a year, every year, and rerun it. What does one more buy you?
The move. Replace an impression with an arithmetic estimate, then test what one variable changes.
Why it is built this way. The one second guess is captured before the arithmetic on purpose, because once you have built the real number you can no longer remember what you assumed, and the gap between the two is the entire exercise. Counting month by month rather than estimating a rate is deliberate, since a remembered rate is a feeling and a counted one is not. The final step exists because a number like this can leave you with nothing but dread, and dread is not useful. Adding one occasion a year is the only variable in the calculation you actually control.
What it does not show. Nothing here says you will act on the number, and nothing says people who do this see anyone more often. This is also not a prediction about anyone’s life expectancy, and it should not be treated as one. It is an estimate built from your own inputs, and its only claim on you is that you built it.

5. The Good Clothes

Go and get the thing you never use because you are saving it: the jacket bought for a life you were expecting, the good shoes, the pen, the notebook too nice to write in. Use it right now on the most ordinary thing it could possibly do, the jacket on while you wash a plate, tomorrow’s list in the notebook. Then say out loud what you were saving it for, the actual occasion you had in mind, and whether that occasion still exists in your life. Rule on it: a real date in the diary, ordinary use starting this week, or out.
The move. Act first, then examine the belief the action exposed.
Why it is built this way. The order is the design. If you ask the question first, you get a reasonable sounding answer about wanting to look after nice things. If you make somebody wear the good jacket to wash a plate first, the question lands on a person who has just felt something, usually a small resistance they were not expecting. Saying the occasion out loud matters, because most of these occasions turn out to belong to a version of your life that either already happened or is not coming. The three way ruling exists so the question cannot be enjoyed and then dropped, and one of the three options is deliberately to keep saving it, as long as you name a date.
What it does not show. No claim about your relationship with money or objects, and nothing here is a diagnosis of anything. We think most people are saving things for a life they are no longer living. We have not tested that, and neither has anyone else that we can find.

6. The Traditions You Already Keep

List the things you or your household already do on repeat without ever having decided to: the meal that appears every week, the mug that is yours, the way birthdays start, the goodbye you always use. Get five down; they exist, you have just never called them anything. Now pick the one an outsider would mistake for a tradition. Write its rule as if it were official: when it happens, what must happen, who does what. Then the harder choice: if you could keep only one of the five for the next fifty years, which survives? And which one would you quietly let die? Say what the kept one preserves.
The move. Put a name and a rule on a structure you already live inside, then decide what it is for.
Why it is built this way. The premise is that a tradition is not invented, it is noticed: repetition plus nobody ever deciding. “They exist, you have just never called them anything” is the reframe the whole question rests on, and five is the forcing number, because one repeated thing is an anecdote and five is how your household actually runs. Writing the rule as if it were official is there to close the gap between a habit and an institution. A habit survives on vagueness, and the moment you have to say when it happens, what must happen and who does what, you find out whether there is anything under the fondness. The fifty year fork is the hard part on purpose. You cannot pick the survivor without saying what each one is for, and “which would you quietly let die” makes you admit that some of what repeats in your life repeats out of momentum rather than meaning. The kept one has to preserve something you can name, or you are keeping a shell.
What it does not show. Nothing says naming these changes them, your household never has to adopt the rule you wrote, and nothing says the one you keep survives fifty years. None of that has been tested. The list is the excuse. Deciding what the kept one protects is the exercise.

7. The Wear Marks

Before you get up, write down the three objects you think carry the most wear from your own hands. Now walk your home and find the real ones: the shine, the dent, the frayed patch, the worn key. Not the oldest things, the most used. Then find the one thing sitting untouched that you bought because you were going to use it. Put your three guesses next to what you actually found. Write two sentences the wear alone would prove: one about where your hours really go, one about who you were planning to be. Then move the untouched thing to where your hands already go, and use it once tonight.
The move. Predict, then check against physical evidence you left without meaning to.
Why it is built this way. Guessing before walking is what gives the walk stakes: without the guess this is a pleasant tour of your home, and with it something is at risk. Wear is a better witness than memory: it cannot be edited afterwards and it accumulated while you were not paying attention. The untouched object is the counterpart of the worn ones, because the gap between what you use and what you bought to use is where intention and behaviour come apart. The last instruction exists because reading evidence and doing nothing is still a tour. Moving the thing to where your hands already go is the smallest act the evidence points at, and one use tonight is the difference between a plan and a shelf. Closing the week here is deliberate. Day one checked what you believe about your people against your phone’s record, and this checks what you believe about yourself against your home’s: the same operation, pointed at a different witness.
What research supports. Strangers who saw only a person’s bedroom or office, and never met them, formed impressions that partly matched how that person and their friends described them. The match was clearest for openness and conscientiousness, and near zero for agreeableness in offices. In the same studies, observers also read meaning into room features that did not track the occupant at all, and part of the agreement between strangers ran through assumptions about sex and race. A room is a weak signal, not a readout. *(Gosling, Ko, Mannarelli & Morris, 2002, Journal of Personality and Social Psychology. Tier B: supported, narrow.)*
What it does not show. That work measured strangers reading rooms. It did not measure whether you can read your own, which is the thing this question asks. We are not claiming improved self-knowledge or improved social perception, and no claim that the moved object stays in use past tonight. What we are claiming is the evidence itself: the wear was already there, and you put it there.

The closing

Here is what a week of this does not do. It does not raise your IQ. It does not make you measurably better at reasoning in general. Almost nothing does, the evidence for far transfer is weak across the board, and the largest company that said otherwise paid $2 million to stop saying it. If someone sells you a daily practice on the promise of general cognitive improvement, they are selling you something that has not been demonstrated and that they would struggle to defend if anyone asked them to.
Here is what it does do, and this part we are confident about because it is a description rather than a prediction. Seven days from now you will have written seven things you would not otherwise have written. Each one came from a question you did not choose, which is the only reason they went anywhere unfamiliar.
That is the whole offer. One question a day that you did not pick, answered in your own words, and at the end of the month what you wrote comes back to you printed on paper.
If that sounds worth fifteen dollars, it is at usethriv.com/whetstone. If it sounds like something you would rather do yourself with a notes app and an alarm, do that. The questions are above.
There is also a free question here every day, a different one from the seven above. Today’s is at whetstone.usethriv.com/daily

Where every claim above comes from

Every source, and what each one does not reach. Two of the seven questions cite research; five cite nothing, because nothing studies them.
Opening. FTC settled with Lumos Labs, $2m paid, $50m entered and suspended; RCT standard for school/work/athletics and age-related decline claims
Source. FTC press release and order, 5 Jan 2016
What it does not reach. The order binds Lumos Labs, not us. A separate, lower standard covers other efficacy claims
Opening. More than 70 scientists, software games do not carry into everyday performance
Source. Stanford Center on Longevity / Max Planck consensus, Oct 2014
What it does not reach. It is about commercial software. It says nothing for or against anything else
Opening. Training on a task improves that task, with little evidence of carryover
Source. Simons et al. (2016), *Psychological Science in the Public Interest* 17(3), 103-186
What it does not reach. The same standard applies to us, and we do not meet it
One study. 319 knowledge workers, 936 tasks, confidence in AI vs confidence in self
Source. Lee, Tankelevitch et al. (2025), CHI
What it does not reach. Self-reported, correlational, one-time survey. No practice was tested
What is built. Material you generate is remembered better than material you read
Source. Generation effect, Slamecka & Graf (1978) and subsequent meta-analysis
What it does not reach. Memory for the generated material only. Says nothing about thinking in general
Rep 1. None
Source. None
What it does not reach. Nothing studies this. Design argument only
Rep 2. None
Source. None
What it does not reach. Nothing studies this. Design argument only
Rep 3. None
Source. None
What it does not reach. Nothing studies this. Design argument only
Rep 4. None
Source. None
What it does not reach. Nothing studies this. Design argument only
Rep 5. None
Source. None
What it does not reach. Tier D admission printed in the rep itself
Rep 6. None
Source. None
What it does not reach. Nothing studies this. Design argument only
Rep 7. Strangers reading a room form partly accurate impressions, and also read meaning into features that carry none
Source. Gosling, Ko, Mannarelli & Morris (2002), *JPSP*
What it does not reach. Does not measure whether you can read your own room