David T Phung
X feed control · the game plan

The steerable feed.

A plan to turn the For You feed from an opaque optimizer into a steerable, self-explaining control system. Stated intent and revealed taste both become first-class ranking signals, every post can explain itself, and the only thing allowed to override you is a narrow, logged public-safety interrupt. Grok is the translator between human language and the ranking stack.

Mapped onto the open-source x-algorithm (home-mixer, the heavy ranker, SimClusters, the visibility library)

The mental model

A feed is a control system, not a settings panel

Every failed slider prototype made the same mistake: it treated a control problem as a configuration problem. A settings panel assumes you know what you want in advance and that it stays stable. Feeds violate both. So model the feed as a closed loop with seven parts, and take a position on each.

Inputs

Signals + intent

Follows, dwell, mutes, snoozes, clicks, and typed or spoken intent treated as a first-class input.

Preference state

The target

Three tiers: declared, inferred, contextual. Each enters a different pipeline stage.

Ranking layer

Condition, don't filter

Candidate gen, light and heavy rank, re-rank. Preferences condition the score, never bolt on after it.

Override layer

Circuit breaker

The only thing that can outvote you. Narrow, labeled, reversible, rate-limited, logged.

Safety layer

Integrity floor

Binds everyone, including overrides and your own preferences. Never amplify harm.

Feedback loop

Correct in place

Every why-this and less/more is a labeled training event captured at the point of consumption.

Learning loop

Become the model

Corrections retrain the ranker, with holdouts guarding against the loop narrowing you over time.

Output

A feed you can talk to

You say what you want, it shows what it did and why, and it tells the truth when it is unsure.

The three open questions, answered

Toggles, ingestion, override

The brief left three questions unresolved. Here is the position on each, in one breath, before the mechanisms.

A. Can preferences be expressed through toggles?

Partly, and the part they can't reach is the part that matters. Toggles carry the legible 20% (topics, languages, mute, snooze, a "less politics" dial). They will never carry taste, which is high-dimensional, contextual, and revealed not stated. The bridge across that gap is Grok compiling natural-language intent into ranking changes. That is why Grok is load-bearing, not decorative.

B. Can the system ingest and actually deliver?

Only if each preference is placed at the stage that matches its kind. A topic mute is a filter. "More long-form" is a re-ranking objective. "I trust these people" is a candidate-source prior. Preference features fail when they are all implemented as the same thing (a filter) at the same place (after ranking), where they fight the optimizer and produce empty feeds. The fix is placement.

C. Should the system ever override a preference?

Yes, exactly once, and it is a scalpel. Withholding a verified imminent-danger alert because someone muted "news" is a worse failure than the interrupt. So the override is real, but engineered as a circuit breaker: narrow, labeled, explained, reversible, rate-limited, and logged. The moment that channel carries anything but your own safety, the product loses trust.

Thesis

The For You feed becomes a steerable, self-explaining control system: you tell it what you want in plain language and by example, it shows you why you are seeing each post, and the only thing that can override your stated preference is a narrow, logged public-safety interrupt, with Grok as the translation layer between human intent and the ranking pipeline.

Reading guide

Every recommendation has a home in the pipeline

The reason this is a systems plan and not a wish list: each mechanism is tagged with where it lives in the stack. These are the six stages, color-coded, used throughout.

candidate gen feature extraction scoring re-ranking policy filtering post-ranking override

The single most important architectural call: preferences condition the ranker, they are not a filter stapled on after it. A filter after ranking fights the optimizer and produces the brittle, empty feeds that killed the slider prototypes. A feature inside the ranker lets the model learn to honor intent while keeping the feed full and alive.

The five product bets

What compounds over 12 months

Five bets, chosen because each makes the next cheaper and the loop tighter. They map one-to-one onto the open questions and the goal of making "Twitter" the wrong word by year end. Tap any bet to open it.

Bet 1

"Why this post" legibility layer

The wedge. Grok-written, signal-grounded explanation on every post, with one-tap less/more.

User problem

People cannot see why a post is in front of them, so no control feels real and the feed feels like something done to them.

Why now

The stack is open. We can attribute a post's slot across the stages that set it (light rank, the heavy ranker's named engagement heads, re-rank), reading the live serving weights, and Grok turns that into one warm, specific sentence.

Why it matters

Legibility is the trust wedge, the cheapest to ship (no retrain to start), and it manufactures the labeled-feedback flywheel every other bet feeds on.

How it compounds

Every less/more is a labeled training event that improves the taste model and the steering compiler. Legibility makes the data the system learns from.

What could fail

Explanations degrade into platitudes ("because you follow X"). Ungrounded, they erode trust instead of building it.

Success looks like

"This explains it" above 70%, less/more used in 1 of every 5 sessions, and a positive feed-satisfaction delta versus the holdback.

explanation surfacecited signalsattributionless/more write-back
Bet 2

Steering, not settings

Natural-language feed control. Type intent, Grok compiles it to a preference vector, applied at the right stages, with live preview. The answer to question A.

User problem

Users feel what they want ("less of this energy, more of that") but cannot express it through a wall of toggles.

Why now

Grok now parses fuzzy intent into a structured, typed preference object reliably. Language captures the direction of a shift, behavior supplies the coordinates: the box edits the gradient, not the embedding.

Why it matters

The killer feature and the literal answer to "can preferences be expressed?" Yes, in language, compiled by Grok, applied where each preference belongs.

How it compounds

The compiled vector is reused everywhere: conditions the heavy ranker, seeds Living Timelines, becomes part of the editable taste model.

What could fail

The compiler mistranslates, or applies intent as a blunt filter that empties the feed. Live preview and one-tap undo are non-negotiable.

Success looks like

Preference-alignment (cosine of asked-direction vs realized shift) up at least +0.3 within 3 sessions, undo under 10%, steering reused not set once.

intent vectorsource reweightconstrained objective
Bet 3

The editable taste model

Inference made legible. "Grok thinks you're into A, B, C, and read more long-form than average. Edit this."

User problem

The system infers a model of you, shapes everything you see with it, and never shows it. You cannot correct what you cannot see.

Why now

User-embedding summarization only now reads as legible interests rather than a creepy dossier. A year ago it was too coarse to use or too raw to show.

Why it matters

The privacy-and-trust differentiator. It defuses the central fear about feeds: that they decide who you are and trap you there.

How it compounds

Edits are gold-standard labels. "No, I am not into crypto" is worth thousands of passive signals, straight into candidate gen and scoring.

What could fail

The profile reads as creepy or wrong. The rule that keeps it a mirror: top-N readable clusters, each traced to K+ signals, editable, never raw embeddings or per-post logs.

Success looks like

25%+ of the power cohort open the taste model in week one, 15%+ edit an interest, and a positive post-edit satisfaction delta.

user-controllable embedding adjunct
Bet 4

Living Timelines

Custom Timelines evolve into named, shareable, Grok-curated lenses that update themselves and carry their own steering. The differentiation bet.

User problem

One feed cannot be everything. "Deep work AI research" and "Sunday sports and friends" are different feeds; forcing them into one stream serves neither.

Why now

Two clocks line up: Custom Timelines proved the demand, and Grok-maintained lenses only now stay fresh autonomously, so a curated feed no longer rots when you stop tending it.

Why it matters

A followed lens becomes a candidate source carrying its curator's authority weight, so curation enters candidate gen as a high-precision source one feed cannot provide. The differentiation bet.

How it compounds

Lenses are a distribution and creator primitive. A great lens can be followed, turning curation into a network product and a high-signal candidate source.

What could fail

Fragmentation. If switching is heavy, people stay in one and it dies. Lens-switching has to be one gesture; the default feed stays excellent.

Success looks like

A median of 2+ active lenses per engaged user, lenses followed across users, rising consumption inside steered lenses without cannibalizing the default.

per-lens candidate sourceper-lens objective
Bet 5

The public-safety interrupt

The override, productized. A narrow, auditable layer that can pierce "I muted news" for verified, time-sensitive, local safety info. The answer to question C.

User problem

Someone who muted "news" still needs the earthquake, the evacuation order, the active-shooter alert in their city. Honoring the mute can get them hurt.

Why now

Grok can corroborate and classify a candidate event across sources and draft a clear, honest label, with calibration, not vibes.

Why it matters

Done right it is a trust asset: proof the system puts your safety above its own engagement and its own consistency. Done wrong, the most dangerous thing here.

How it compounds

The credibility capstone. A platform trusted to interrupt you only for your safety earns the right to steer everything else.

What could fail

A false alert, or any perception the channel carries commercial content. Either one detonates the entire thesis.

Success looks like

High override precision, low regret and reversal, a published transparency report, and users who trust the interrupt rather than fear it.

audited injection above the foldintegrity floor

The preference model

Four buckets, each routed to a different stage

The discipline is putting each preference in the right bucket, then routing it to the right stage. Mixing them up is the original sin of the slider era.

Direct control

Declared, stable, legible: topics on/off, languages, a "less politics" dial, mute, block, snooze, sensitive-content settings, trusted lists, per-timeline steering text.

candidate genpolicy filteringsoft dials

Implicit

Revealed, never asked, used as features: dwell, completion, scroll-back, profile clicks, share, bookmark, who you DM. Never shown as toggles, because asking would corrupt them.

feature extractionscoring

Inferred

Constructed, shown for correction: your taste vector, long-form vs quick-hit propensity, out-of-network tolerance, per-account weights. Surfaced through the editable taste model.

feature extraction, exposed

Overrideable

Narrow, the user's benefit only: integrity floors, then imminent physical danger, then high-confidence civic info. Everything else is not overrideable.

post-ranking override

The governing rule, enforced everywhere: the system may surface what you asked to hide for your own safety. It may never override what you asked to block. And it may never inject commercial or engagement content through the override channel. Overrides exist for the user, never for the platform.

The override decision framework · interactive

When the feed is allowed to outvote you

The override fires only when every hard gate passes and a loss-minimizing comparison says firing beats holding. Gates decide whether allowed; the comparison decides whether worth it. Pre-vetted authoritative sources fire on a fast path; novel events wait for human sign-off and default to not firing if that queue saturates. Move the sliders, or load a scenario, and watch the policy decide.

0.00
override score

HOLD
Grok proposes, policy disposes
Adjust the factors to see how the policy reasons.

OverrideScore = Severity · Confidence · Locality · Timeliness · (1 − Redundancy), every factor in [0,1]. Multiplicative here to show the intuition: any near-zero factor kills the override on its own, so an alert for a city you are not in never fires no matter how severe. The live policy is stricter than this slider. Confidence and Locality are hard gates, not soft factors. The decision weighs the harm of a wrong alert against the harm of a missed one, not just the upside. And Tier 1 requires a designated-authoritative source with provably independent corroboration, because the primary attack is sockpuppets faking agreement to spoof a local emergency.

How ranking reads the signals

Each signal: meaning, failure mode, stage, trust

The same signal can mean opposite things. Dwell is interest, or it is outrage. The model has to know the difference, so every signal carries a failure mode and a stage.

SignalWhat it meansFailure modeStageTrust
Topic toggleDeclarative interest / disinterestTopics are fuzzy; keyword match missescand gen re-rankHigh precision, low recall
Explicit followStrong stated interest in a sourceStale follows you forgotcand genHigh, decay by recency
Mute / snoozeStop showing thisOver-muting collapses recallfilter / re-rankVery high, snooze must decay
Dwell timeRevealed attentionConfounded by confusion and outragefeatures scoringMedium, use good dwell only
Engagement qualityHow meaningful the interaction wasLikes are cheap, replies can be hostilescoringHigh when quality-weighted
Negative feedbackThe strongest steering signalAsymmetric; one "show less" outweighs many passivescoring + taste modelHighest per event
Recency + contextFreshness and situationOver-indexing floods you during spikesdecay re-rankContext-dependent

Two cross-cutting rules. Asymmetry: an explicit negative ("show less", mute) is worth far more than a passive positive (a like), because users spend negatives deliberately. Good dwell over raw dwell: dwell ending in a positive action is interest; dwell ending in a mute or fast scroll-away is often outrage, and rewarding it builds a rage feed.

Mapping onto x-algorithm · the answer to B

Where every mechanism lands in the real stack

Grounded in the open-source services. Each product mechanism has a concrete home, which is the difference between "express a preference" and "deliver on it".

Product mechanismx-algorithm homeStage
Trusted people, follow prefsEarlybird + RealGraph (in-network)candidate gen
"What my circle is into"UTEG / DirectUtegcandidate gen
Topic toggles (semantic)SimClusters communities, reweightedcandidate gen
Niche discovery, "more like this"TweetMixer / TwHIN + ContentExplorationcandidate gen
Intent vector + taste modelrecap (heavy ranker) input, served by navifeature extraction
Preference alignmentnew head beside the ~17 engagement headsscoring
Soft dials + steering objectivesre-ranker constraintsre-ranking
Mutes, blocks, integrity floorvisibility / trust-and-safety librarypolicy filtering
Public-safety interruptfinal injection above the foldpost-ranking override

How Grok changes the product

Five jobs only a strong model can do

Grok is load-bearing in five distinct jobs. That is what makes the layer defensible, and it is the difference between a feed you configure and a feed you can talk to.

  • Preference interpretation. Natural-language intent in, structured preference vector out, routed to the right stages. The irreplaceable job. fx cg
  • Semantic topic understanding. Replace brittle keyword topics so "less politics" actually catches political content and "more AI research" separates research from hype. cg rr
  • Summarization and explanation. The why-this sentence, the "what you missed" digest, the override label. ov
  • Uncertainty handling. Calibrated uncertainty. On low confidence it asks instead of assuming, and abstains on ambiguous override candidates. This keeps it from confidently steering you wrong.
  • Safety and override support. Grok corroborates an event across sources, classifies its tier, and drafts the label, but never authorizes Tier 1 alone. Grok proposes, the audited policy layer disposes. ov

The interface model

One knob visible, infinite depth available

The slider graveyard failed because it front-loaded complexity. Invert it: progressive disclosure in three layers, where 95% of users never see a control panel.

Layer 0

Everyone, invisible

The feed just gets better. Long-press any post for "why this?" Inline "show less / more like this." No new surface to learn.

Layer 1

One tap

A single "Tune" affordance: one text-or-voice box ("tell the feed what you want") plus a few smart chips, with a live preview and one-tap undo. The box is the interface; Grok does the compiling.

Layer 2

Power users

The full steering panel, the editable taste model, the Living Timelines manager, override settings, and the transparency log. Deep, but opt-in.

Show the result, not the machinery. Every control has an immediately visible effect (the preview, the rank drop, the explanation), so adjustment feels like steering, not configuring.

Rollout + instrumentation

Ship the trust, then the steering

Sequence.

  • MVP: why-this + inline less/more to a 1% holdback. No retrain to start.
  • Power-user: Living Timelines, editable taste model, override in shadow mode with a red team.
  • Mainstream: default the Tune affordance; ship the interrupt only after a long shadow trial proves precision, one tier, one geo.

Metrics, not raw engagement.

  • Steerability (primary): did the feed move in the direction asked.
  • Legibility: "this explains it" rate.
  • Trust: survey, override regret, reversal, mute trend.
  • Healthy engagement: good dwell, reply-engaged-back, day-N return.
  • Guardrails: diversity index, integrity, feedback-loop-collapse monitor.

Experiment plan: shadow mode before anything ships (especially overrides), long-horizon holdouts to catch feedback-loop collapse that short A/Bs miss, switchback designs for live events and overrides, and a standing override red team. The override never leaves shadow mode without a passing precision bar and a kill switch.

Brutally honest risk assessment

Where this breaks, and the fix

Too complex

The steering surface metastasizes into a settings maze. Fix: Layer-1 default is one box plus a preview; Grok compiles; Layer 2 is opt-in. If a normal user must learn a control panel, we failed.

Too shallow

"Why this" decays into "because you follow X". Fix: ground explanations in real per-post feature attributions. A fake explanation is worse than none.

Manipulable

Actors game "more like this" to farm reach, or trigger the override channel. Fix: authorization gate, corroboration, rate limits, and commercial content never rides the override. The integrity floor binds even the user.

Hurts trust

A wrong override or any whiff of manipulation detonates the thesis in one incident. Fix: conservative thresholds, corroboration over speed, labeling, reversibility, a public log. Tuned to under-fire, not over-fire.

Execution roadmap

30, 90, 180, end of year

Next 30 days
  • Why-this + less/more to a 1% holdback
  • Stand up the preference-vector schema
  • Prototype the Grok intent compiler offline
  • Draft override taxonomy, begin shadow logging
Next 90 days
  • NL steering box to power users behind a flag, live preview
  • Editable taste model v1
  • Override in shadow mode, red team engaged
  • Steerability / legibility / trust dashboard
Next 180 days
  • Living Timelines to GA-candidate
  • Steering to a broader cohort; mainstream Tune test
  • Safety interrupt limited live pilot: one tier, one geo
  • Transparency report live
End of year
  • Feed steerable in plain language
  • Every post can explain itself
  • Multiple living timelines per user
  • A narrow, audited interrupt behind a public log. "Twitter" is the wrong word.

What I would ship first

"Why this post" + inline less/more, to a 1% holdback, this month. Lowest risk, highest trust: it rides on the existing pipeline, reads the live serving weights, needs no retrain to begin, and starts manufacturing the labeled data every other bet depends on. Legibility earns the right to steer.

What I would not build

  • A wall of per-topic percentage sliders.
  • "Pure chronological" as the headline answer (it stays as an escape hatch).
  • Any path for ads or engagement to ride the override channel.
  • A fully autonomous override with no human gate on Tier 1.
  • Opaque taste inference with no edit path.
  • Gamified preference scores or streaks.

The full document

Read the complete game plan

This page is the skimmable version. The full master plan carries all 16 sections plus four appendices: a one-page PRD, the model and policy spec, the experimentation matrix, and the engineering milestone plan by sprint.

Open the master plan (markdown)