Pre-1.0. Not intended for public use yet. The API is subject to undocumented change, and these pages may already be wrong about it.

Skip to content

Storage Architecture ​

Introduction ​

Storyfeed stores each activity once, in feed_activities, and writes the supporting rows used to retrieve it efficiently at publication: the entities' labels, the groups the activity can join, and an index of who and what it involves. Retrieving a feed queries those rows. Storyfeed does not use Laravel's cache for feed data.

This page follows one publish into the database and one page of the feed back out. Schema has the diagram of the tables and every column.

Writing an Activity ​

A customer places an order:

Where the order is placed: a controller, an action, a listener
php
use Storyfeed\Facades\Storyfeed;

Storyfeed::activity()
    ->by($customer)
    ->action('place', $order)
    ->to($shop)
    ->publish();

publish() stamps published_at and runs the verb's story middleware. The innermost step stores the activity in one database transaction:

  1. Snapshots. Each Feedable entity's toFeed() output is upserted into feed_snapshots, one row per entity, and the activity's cached_{role}_id columns are set to those rows. By default, the feed gets labels from these snapshots without loading your models. Presentation resolvers can explicitly call FeedContext::model() to hydrate live models, relations, or counts, adding those queries to the retrieval cost.
  2. The activity. One row in feed_activities, with a new ULID uid. The uid becomes the payload's id.
  3. Groupings. One feed_groupings row per axis the activity has a key for, written in a single insert. The hash is the axis key: the fields the axis compares, joined, ending with the day (or the verb's grouping period).
  4. Participants. One feed_participants row per filled role, with published_at copied from the activity.
  5. Curation. For each group the activity joined, Storyfeed decides which one it appears under in live() and sets winner on that row.

Storyfeed then dispatches ActivityPublished, after the outermost transaction commits. On the way back out of the middleware, the batch middleware adds the activity to the actor's current batch in a transaction of its own.

The Rows One Publish Writes ​

For the order above, with a customer, an order and a shop, all Feedable, and the default axes and middleware:

TableRowsNotes
feed_activities1
feed_snapshots3 upsertedshared: the customer's row serves every activity that names the customer
feed_groupings9actors, targets, object, repeat; summary.hour, summary.day, summary.week, summary.month; batch
feed_participants3actor, object, target
feed_batches0 or 1a new row only when the customer has no open batch; otherwise one update
feed_batch_locks1 upsertedone per customer, reused

An activity with fewer roles writes fewer rows. An activity with no actor has no summary.* rows and joins no batch.

Denormalized Columns ​

These copies avoid repeated lookups and computations during retrieval.

CopyPurpose during retrieval
cached_{role}_idresolving labels from your models; one whereIn per role on feed_snapshots
feed_groupings.hashcomputing group keys over history; a group is every row sharing (bucket, hash)
feed_groupings.winnerdeciding each activity's group on every retrieval
feed_participantsone entity lookup instead of an OR across role pairs; individual role indexes can serve OR branches, but the plan and ordering cost depend on the database planner
feed_participants.published_atcarries activity time in the entity index; involving() selects matching activity IDs here, while the outer activity query orders the results

Retrieving a Page ​

A controller, or wherever the feed is retrieved
php
use Storyfeed\Facades\Storyfeed;

Storyfeed::feed()->get();

Grouping is decided when activities are written. Retrieval selects groups that already exist. A page of a live() feed takes two phases.

Phase one selects the page. Two logical streams return feed items, newest first; they can require more than two SQL queries:

  • The group stream joins each published activity to its winning feed_groupings row and groups by (bucket, hash). Each group's latest published_at places it in the feed. Storyfeed runs this aggregate over the newest 16 × limit activities first, then 256 × limit, and finally the whole eligible history. Each bounded attempt probes a window floor and runs the aggregate. It stops widening when it finds more than limit candidates (the limit + 1 lookahead), or reaches an unbounded attempt. A windowed result also recounts its selected groups across eligible history.
  • The solo stream returns activities that have no winning grouping row.

Storyfeed merges the two, takes limit items (30 by default), and encodes the position of the last one as next_cursor.

Phase two fills in the selected groups. For the groups on this page only, Storyfeed fetches up to grouping.children_limit members each (25 by default), counts distinct entities per role, and eager-loads the members' snapshots. A group with one member is shown as a single activity.

From Three Orders to One Row ​

The customer places three orders, minutes apart. Each order writes the rows in the table above. In log(), they are three rows:

Their repeat rows share one hash: same actor, same verb, same type of object, same shop, same day. No other axis qualifies (one actor, one shop, three different orders), so curation sets winner on repeat for each. The group stream returns one group with count: 3, and its headline comes from the verb's repeat headline in routes/feed.php:

ES
Erica Sinclair placed 3 orders with Scoops Ahoy

Curation checks the axes in order: actors (3 different actors by default), targets, object, then repeat when none qualifies. The thresholds are in Aggregation.

Feed Modes ​

ModeRetrieves
live()the winner grouping row of each activity, or repeat when none is stamped
summary()the summary.{period} row: one group per actor per period
log()atomic activities without aggregation; one main SELECT checks composite claims in feed_groupings with NOT EXISTS to suppress parent stories, plus snapshot loads

Pagination ​

Cursors are keyset positions, not offsets. A log() cursor holds the last row's published_at and id; a grouped cursor also holds the group's axis and hash. A page deep in the feed costs what the first page costs in log(). Treat the cursor as opaque, and use it only with the mode that produced it.

Keeping Stored Rows Current ​

Storyfeed keeps its copies current by writing them again. These are the rules:

CopyRewritten when
a snapshotthe model is saved, the model appears in a publish, storyfeed:trickle finds it stale, storyfeed:rebuild runs, or php artisan optimize invokes storyfeed:cache-snapshots for a bounded recent refresh (skipped when the database is unavailable)
a stale snapshottoFeed() changes shape: each snapshot stores a shape fingerprint, and the trickle refreshes rows whose fingerprint differs
winneran activity is published into the group, or deleted from it; scheduled storyfeed:curate runs hourly, with a configurable lookback
hashstoryfeed:curate --rehash
role columns, snapshots, participants, groupingsa model is deleted: its activities are repointed to a tombstone, in chunks of 500
a batch's closed_atthe actor's next publish arrives after closes_at, or storyfeed:close-batches runs

Scheduled curation defaults to the last two days (storyfeed.curate.window). Weekly and monthly verb grouping widens that window for the affected activities. A null or nonpositive configured window removes the time bound.

When both the stored and incoming source timestamps are known, a snapshot update rejects an older model timestamp. Without a usable incoming timestamp, the update is accepted and clears the watermark. A missing stored timestamp also prevents an ordering comparison.

When stored history changes in a way a client cannot reconcile, Storyfeed writes a new sync_token to feed_meta. Every page carries it. These write one:

ChangeWhen
storyfeed:bundle, storyfeed:curate --rehash, storyfeed:healeach run that rewrites history
deleting or restoring a modelwhen any activity is repointed
releasing a compositewhen a composite is released
pruning or purgingwhen a group loses members

See Handling a Changed Feed.

storyfeed:cache caches your definitions from routes/feed.php, not feed data. See Caching Definitions.

Costs at Scale ​

The default example above writes 1 activity row, 9 grouping rows, and 3 participant rows. Nine is not a ceiling: custom axes can add or replace keys, and each emitted hash produces a grouping row. Table sizes depend on your roles, axes, entities, and retention policy. feed_snapshots grows with Feedable entities, including those saved without publishing an activity.

The table lists indexes available to these query shapes. Plans were checked with 50,000 activities on MariaDB 10.11, PostgreSQL 18, and SQLite 3.45; MySQL 8.x was not tested. The planner can choose a different index or a scan as data distribution and query constraints change. Schema lists every index.

QueryAvailable Index
log(), and the solo streamfeed_activities (published_at, id)
->actor(), ->object(), ->target(), ->context()feed_activities ({role}_type, {role}_id, published_at, id)
->involving()feed_participants (entity_type, entity_id, published_at, activity_id)
each activity's winning row in live()feed_groupings (activity_id)
each activity's period row in summary()feed_groupings unique (activity_id, bucket)
a group's membersfeed_groupings (bucket, hash)
the solo stream's repeat and composite checks, per activityfeed_groupings unique (activity_id, bucket)

Two indexes serve work other than feed retrieval. Publishing finds an actor's open batches through the feed_batch_locks primary key, (actor_type, actor_id). The aggregates check in storyfeed:doctor uses feed_groupings (winner, bucket, hash).

The solo stream's winner check can scan much of the stored history even when it returns no items. In the measured fixture, it dominated live() retrieval cost on MariaDB and PostgreSQL. Composite checks also varied: some plans used (bucket, hash) instead of the unique (activity_id, bucket) index. Treat the table as available access paths, not a promise that every listed index is chosen for every request.

Retention removes old activities with their grouping and participant rows.

Comparison With a Single Log Table ​

A single table with one row per event can render log(). The questions a feed asks next are the ones a single table cannot answer from an index:

QuestionOne tableStoryfeed
"What should this row say?"load each model, per rowthe snapshot, eager-loaded
"Placed 4 orders"group by expressions over history, on every retrievalrows already share a hash
"Which group does this activity belong to?"decided on every retrievaldecided once, at publish
"Everything involving this order"OR across every role columnone indexed lookup
"The order was deleted"the label is gonea tombstone keeps the story readable

The extra rows are written once, when the activity is published. Subsequent retrieval uses them.

Released under the MIT License.