A plan to turn the For You feed from an opaque optimizer into a steerable, self-explaining control system. Stated intent and revealed taste both become first-class ranking signals, every post can explain itself, and the only thing allowed to override you is a narrow, logged public-safety interrupt. Grok is the translator between human language and the ranking stack.
The mental model
Every failed slider prototype made the same mistake: it treated a control problem as a configuration problem. A settings panel assumes you know what you want in advance and that it stays stable. Feeds violate both. So model the feed as a closed loop with seven parts, and take a position on each.
Follows, dwell, mutes, snoozes, clicks, and typed or spoken intent treated as a first-class input.
Three tiers: declared, inferred, contextual. Each enters a different pipeline stage.
Candidate gen, light and heavy rank, re-rank. Preferences condition the score, never bolt on after it.
The only thing that can outvote you. Narrow, labeled, reversible, rate-limited, logged.
Binds everyone, including overrides and your own preferences. Never amplify harm.
Every why-this and less/more is a labeled training event captured at the point of consumption.
Corrections retrain the ranker, with holdouts guarding against the loop narrowing you over time.
You say what you want, it shows what it did and why, and it tells the truth when it is unsure.
The three open questions, answered
The brief left three questions unresolved. Here is the position on each, in one breath, before the mechanisms.
A. Can preferences be expressed through toggles?
Partly, and the part they can't reach is the part that matters. Toggles carry the legible 20% (topics, languages, mute, snooze, a "less politics" dial). They will never carry taste, which is high-dimensional, contextual, and revealed not stated. The bridge across that gap is Grok compiling natural-language intent into ranking changes. That is why Grok is load-bearing, not decorative.
B. Can the system ingest and actually deliver?
Only if each preference is placed at the stage that matches its kind. A topic mute is a filter. "More long-form" is a re-ranking objective. "I trust these people" is a candidate-source prior. Preference features fail when they are all implemented as the same thing (a filter) at the same place (after ranking), where they fight the optimizer and produce empty feeds. The fix is placement.
C. Should the system ever override a preference?
Yes, exactly once, and it is a scalpel. Withholding a verified imminent-danger alert because someone muted "news" is a worse failure than the interrupt. So the override is real, but engineered as a circuit breaker: narrow, labeled, explained, reversible, rate-limited, and logged. The moment that channel carries anything but your own safety, the product loses trust.
Thesis
Reading guide
The reason this is a systems plan and not a wish list: each mechanism is tagged with where it lives in the stack. These are the six stages, color-coded, used throughout.
The single most important architectural call: preferences condition the ranker, they are not a filter stapled on after it. A filter after ranking fights the optimizer and produces the brittle, empty feeds that killed the slider prototypes. A feature inside the ranker lets the model learn to honor intent while keeping the feed full and alive.
The five product bets
Five bets, chosen because each makes the next cheaper and the loop tighter. They map one-to-one onto the open questions and the goal of making "Twitter" the wrong word by year end. Tap any bet to open it.
"Why this post" legibility layer
The wedge. Grok-written, signal-grounded explanation on every post, with one-tap less/more.
▾User problem
People cannot see why a post is in front of them, so no control feels real and the feed feels like something done to them.
Why now
The stack is open. We can attribute a post's slot across the stages that set it (light rank, the heavy ranker's named engagement heads, re-rank), reading the live serving weights, and Grok turns that into one warm, specific sentence.
Why it matters
Legibility is the trust wedge, the cheapest to ship (no retrain to start), and it manufactures the labeled-feedback flywheel every other bet feeds on.
How it compounds
Every less/more is a labeled training event that improves the taste model and the steering compiler. Legibility makes the data the system learns from.
What could fail
Explanations degrade into platitudes ("because you follow X"). Ungrounded, they erode trust instead of building it.
Success looks like
"This explains it" above 70%, less/more used in 1 of every 5 sessions, and a positive feed-satisfaction delta versus the holdback.
Steering, not settings
Natural-language feed control. Type intent, Grok compiles it to a preference vector, applied at the right stages, with live preview. The answer to question A.
▾User problem
Users feel what they want ("less of this energy, more of that") but cannot express it through a wall of toggles.
Why now
Grok now parses fuzzy intent into a structured, typed preference object reliably. Language captures the direction of a shift, behavior supplies the coordinates: the box edits the gradient, not the embedding.
Why it matters
The killer feature and the literal answer to "can preferences be expressed?" Yes, in language, compiled by Grok, applied where each preference belongs.
How it compounds
The compiled vector is reused everywhere: conditions the heavy ranker, seeds Living Timelines, becomes part of the editable taste model.
What could fail
The compiler mistranslates, or applies intent as a blunt filter that empties the feed. Live preview and one-tap undo are non-negotiable.
Success looks like
Preference-alignment (cosine of asked-direction vs realized shift) up at least +0.3 within 3 sessions, undo under 10%, steering reused not set once.
The editable taste model
Inference made legible. "Grok thinks you're into A, B, C, and read more long-form than average. Edit this."
▾User problem
The system infers a model of you, shapes everything you see with it, and never shows it. You cannot correct what you cannot see.
Why now
User-embedding summarization only now reads as legible interests rather than a creepy dossier. A year ago it was too coarse to use or too raw to show.
Why it matters
The privacy-and-trust differentiator. It defuses the central fear about feeds: that they decide who you are and trap you there.
How it compounds
Edits are gold-standard labels. "No, I am not into crypto" is worth thousands of passive signals, straight into candidate gen and scoring.
What could fail
The profile reads as creepy or wrong. The rule that keeps it a mirror: top-N readable clusters, each traced to K+ signals, editable, never raw embeddings or per-post logs.
Success looks like
25%+ of the power cohort open the taste model in week one, 15%+ edit an interest, and a positive post-edit satisfaction delta.
Living Timelines
Custom Timelines evolve into named, shareable, Grok-curated lenses that update themselves and carry their own steering. The differentiation bet.
▾User problem
One feed cannot be everything. "Deep work AI research" and "Sunday sports and friends" are different feeds; forcing them into one stream serves neither.
Why now
Two clocks line up: Custom Timelines proved the demand, and Grok-maintained lenses only now stay fresh autonomously, so a curated feed no longer rots when you stop tending it.
Why it matters
A followed lens becomes a candidate source carrying its curator's authority weight, so curation enters candidate gen as a high-precision source one feed cannot provide. The differentiation bet.
How it compounds
Lenses are a distribution and creator primitive. A great lens can be followed, turning curation into a network product and a high-signal candidate source.
What could fail
Fragmentation. If switching is heavy, people stay in one and it dies. Lens-switching has to be one gesture; the default feed stays excellent.
Success looks like
A median of 2+ active lenses per engaged user, lenses followed across users, rising consumption inside steered lenses without cannibalizing the default.
The public-safety interrupt
The override, productized. A narrow, auditable layer that can pierce "I muted news" for verified, time-sensitive, local safety info. The answer to question C.
▾User problem
Someone who muted "news" still needs the earthquake, the evacuation order, the active-shooter alert in their city. Honoring the mute can get them hurt.
Why now
Grok can corroborate and classify a candidate event across sources and draft a clear, honest label, with calibration, not vibes.
Why it matters
Done right it is a trust asset: proof the system puts your safety above its own engagement and its own consistency. Done wrong, the most dangerous thing here.
How it compounds
The credibility capstone. A platform trusted to interrupt you only for your safety earns the right to steer everything else.
What could fail
A false alert, or any perception the channel carries commercial content. Either one detonates the entire thesis.
Success looks like
High override precision, low regret and reversal, a published transparency report, and users who trust the interrupt rather than fear it.
The preference model
The discipline is putting each preference in the right bucket, then routing it to the right stage. Mixing them up is the original sin of the slider era.
Declared, stable, legible: topics on/off, languages, a "less politics" dial, mute, block, snooze, sensitive-content settings, trusted lists, per-timeline steering text.
Revealed, never asked, used as features: dwell, completion, scroll-back, profile clicks, share, bookmark, who you DM. Never shown as toggles, because asking would corrupt them.
Constructed, shown for correction: your taste vector, long-form vs quick-hit propensity, out-of-network tolerance, per-account weights. Surfaced through the editable taste model.
Narrow, the user's benefit only: integrity floors, then imminent physical danger, then high-confidence civic info. Everything else is not overrideable.
The governing rule, enforced everywhere: the system may surface what you asked to hide for your own safety. It may never override what you asked to block. And it may never inject commercial or engagement content through the override channel. Overrides exist for the user, never for the platform.
The override decision framework · interactive
The override fires only when every hard gate passes and a loss-minimizing comparison says firing beats holding. Gates decide whether allowed; the comparison decides whether worth it. Pre-vetted authoritative sources fire on a fast path; novel events wait for human sign-off and default to not firing if that queue saturates. Move the sliders, or load a scenario, and watch the policy decide.
OverrideScore = Severity · Confidence · Locality · Timeliness · (1 − Redundancy), every factor in [0,1]. Multiplicative here to show the intuition: any near-zero factor kills the override on its own, so an alert for a city you are not in never fires no matter how severe. The live policy is stricter than this slider. Confidence and Locality are hard gates, not soft factors. The decision weighs the harm of a wrong alert against the harm of a missed one, not just the upside. And Tier 1 requires a designated-authoritative source with provably independent corroboration, because the primary attack is sockpuppets faking agreement to spoof a local emergency.
How ranking reads the signals
The same signal can mean opposite things. Dwell is interest, or it is outrage. The model has to know the difference, so every signal carries a failure mode and a stage.
| Signal | What it means | Failure mode | Stage | Trust |
|---|---|---|---|---|
| Topic toggle | Declarative interest / disinterest | Topics are fuzzy; keyword match misses | cand gen re-rank | High precision, low recall |
| Explicit follow | Strong stated interest in a source | Stale follows you forgot | cand gen | High, decay by recency |
| Mute / snooze | Stop showing this | Over-muting collapses recall | filter / re-rank | Very high, snooze must decay |
| Dwell time | Revealed attention | Confounded by confusion and outrage | features scoring | Medium, use good dwell only |
| Engagement quality | How meaningful the interaction was | Likes are cheap, replies can be hostile | scoring | High when quality-weighted |
| Negative feedback | The strongest steering signal | Asymmetric; one "show less" outweighs many passive | scoring + taste model | Highest per event |
| Recency + context | Freshness and situation | Over-indexing floods you during spikes | decay re-rank | Context-dependent |
Two cross-cutting rules. Asymmetry: an explicit negative ("show less", mute) is worth far more than a passive positive (a like), because users spend negatives deliberately. Good dwell over raw dwell: dwell ending in a positive action is interest; dwell ending in a mute or fast scroll-away is often outrage, and rewarding it builds a rage feed.
Mapping onto x-algorithm · the answer to B
Grounded in the open-source services. Each product mechanism has a concrete home, which is the difference between "express a preference" and "deliver on it".
| Product mechanism | x-algorithm home | Stage |
|---|---|---|
| Trusted people, follow prefs | Earlybird + RealGraph (in-network) | candidate gen |
| "What my circle is into" | UTEG / DirectUteg | candidate gen |
| Topic toggles (semantic) | SimClusters communities, reweighted | candidate gen |
| Niche discovery, "more like this" | TweetMixer / TwHIN + ContentExploration | candidate gen |
| Intent vector + taste model | recap (heavy ranker) input, served by navi | feature extraction |
| Preference alignment | new head beside the ~17 engagement heads | scoring |
| Soft dials + steering objectives | re-ranker constraints | re-ranking |
| Mutes, blocks, integrity floor | visibility / trust-and-safety library | policy filtering |
| Public-safety interrupt | final injection above the fold | post-ranking override |
How Grok changes the product
Grok is load-bearing in five distinct jobs. That is what makes the layer defensible, and it is the difference between a feed you configure and a feed you can talk to.
The interface model
The slider graveyard failed because it front-loaded complexity. Invert it: progressive disclosure in three layers, where 95% of users never see a control panel.
The feed just gets better. Long-press any post for "why this?" Inline "show less / more like this." No new surface to learn.
A single "Tune" affordance: one text-or-voice box ("tell the feed what you want") plus a few smart chips, with a live preview and one-tap undo. The box is the interface; Grok does the compiling.
The full steering panel, the editable taste model, the Living Timelines manager, override settings, and the transparency log. Deep, but opt-in.
Show the result, not the machinery. Every control has an immediately visible effect (the preview, the rank drop, the explanation), so adjustment feels like steering, not configuring.
Rollout + instrumentation
Sequence.
Metrics, not raw engagement.
Experiment plan: shadow mode before anything ships (especially overrides), long-horizon holdouts to catch feedback-loop collapse that short A/Bs miss, switchback designs for live events and overrides, and a standing override red team. The override never leaves shadow mode without a passing precision bar and a kill switch.
Brutally honest risk assessment
The steering surface metastasizes into a settings maze. Fix: Layer-1 default is one box plus a preview; Grok compiles; Layer 2 is opt-in. If a normal user must learn a control panel, we failed.
"Why this" decays into "because you follow X". Fix: ground explanations in real per-post feature attributions. A fake explanation is worse than none.
Actors game "more like this" to farm reach, or trigger the override channel. Fix: authorization gate, corroboration, rate limits, and commercial content never rides the override. The integrity floor binds even the user.
A wrong override or any whiff of manipulation detonates the thesis in one incident. Fix: conservative thresholds, corroboration over speed, labeling, reversibility, a public log. Tuned to under-fire, not over-fire.
Execution roadmap
What I would ship first
"Why this post" + inline less/more, to a 1% holdback, this month. Lowest risk, highest trust: it rides on the existing pipeline, reads the live serving weights, needs no retrain to begin, and starts manufacturing the labeled data every other bet depends on. Legibility earns the right to steer.
What I would not build
The full document
This page is the skimmable version. The full master plan carries all 16 sections plus four appendices: a one-page PRD, the model and policy spec, the experimentation matrix, and the engineering milestone plan by sprint.
Open the master plan (markdown)