People are not
the dataset.
Observes, selects and presents. Does not judge.
A temporary editorial field
HEARD is a temporary editorial field of public voices describing life with AI. It presents the best material that remains publicly available and revalidates from the preceding seven days; it is not a representative or momentary sample. It does not diagnose, profile, infer personal traits or make automated decisions about people. It is neither a survey, a news feed, a social-media wall nor a claim about public opinion.
Four bounded source adapters
Bluesky uses eight explicit English-language searches across six topic baskets—work, creativity, learning, everyday use, reliance or attachment, and doubt or friction—through the official API. A dedicated Preview credential creates a short-lived session held only in process memory; credentials and tokens are neither persisted nor logged. Posts with moderation labels, replies, embeds or an opt-out content visibility declaration are excluded. An unknown declaration is excluded too; v1 checks only the fixed official bsky.social PDS and may therefore miss other hosting providers. No profile enrichment, avatars, biographies, follower graphs or account histories are requested.
Mastodon interleaves at most two pages of public hashtag timelines from up to three explicitly configured HTTPS instances and four reviewed hashtags. Requests stop at the shared 24-candidate ceiling and remain inside a 36-request source ceiling. Federated copies are deduplicated by canonical post and account URL; only origins in the same explicit allowlist qualify. This is not global Mastodon search. Only English, standalone public text posts qualify; private, unlisted, sensitive, media, reblogs and replies are excluded. With no configured instances, Mastodon is disabled. API responses can include profile data; unused fields are discarded, not enriched, and profiles or account histories are not read.
Hacker News uses only its official Firebase API: up to 30 recent IDs from each of newstories and askstories. IDs are deduplicated, each story’s title/type is checked first, and comments are traversed only for stories that match transparent AI-topic terms. Reads stop at 24 eligible source expressions or a 72-request source ceiling; depth is at most two and concurrency at most four. It never reads user histories or a third-party search service.
YouTube uses only the official Data API v3. Six topic searches refresh a saved set of at most twelve video IDs no more than once per 24 hours; normal refreshes request one bounded page of top-level public comments from that set. Replies, channel profiles, histories and enrichment are not requested. The saved set contains only video IDs, topic buckets and expiry metadata—never titles, comments or attribution. The adapter is disabled by default and quota or discovery failure fails closed. These routes reach different audiences and have different discovery bias; none is representative. Reddit remains a future adapter.
Expressions, not attributes
Hard eligibility is limited to source validity, publication within the previous seven days and survival through snapshot expiry plus five minutes, public/current status, supported standalone format, spam or explicit commercial promotion, provider moderation and sensitive-material exclusions, suppression, deduplication, and clear relevance to lived experience or opinion about AI. It does not use provider quotas, topic quotas or target field size. One candidate expression per provider-scoped source account is allowed.
After those gates, one bounded OpenAI Structured Outputs request classifies at most thirty candidates as strong, acceptable or weak and assigns only fixed evidence-type, stance, topic, score and evidence-flag vocabularies. Input contains exact text, source type and broad recency—not usernames, handles, DIDs, native IDs, URLs, avatars or profile metadata. Candidate text is treated as untrusted data. The model does not rewrite, quote, summarize or generate a public substitute. Classification is editorial and fallible, not profiling, diagnosis or sentiment analysis.
The publication contract is quality-first: a field needs at least three selected source expressions from three provider-scoped accounts, including at least one strong expression. Every remaining expression may be strong or acceptable, with at most two considered opinions. Weak, promotional, advice or announcement, hypothetical-only and sensitive expressions are excluded. Six to ten voices, multiple providers, stance/topic diversity and no more than three voices per provider where alternatives exist are soft preferences. A valid single-provider field is allowed and labelled narrow.
Seven narrow literal indicators remain optional aggregate observations: replacement concern, reliance expressed, agency retained, curiosity expressed, creative possibility, workflow friction and trust questioned. A category appears only for at least ten distinct source accounts; zero categories is valid. The public field exposes only aggregate signal count, source count, an unambiguous dominant topic, and source-balanced or narrow (source distribution only, never balanced viewpoints or measured public sentiment). Individual model scores, topics, decisions and motifs are never exposed or attached to an opened voice. The field is editorial, not a survey or public-opinion estimate.
Choose to hear a voice
The field contains ephemeral signal IDs, source labels, broad recency, presentation capability and aggregate counts. No post text, source URL, DID, username, channel ID, video ID, comment ID or other native identifier is present in the ordinary field response. The original voice is retrieved only after the visitor selects an individual signal. That one selection opens the panel and revalidates the active source, hard eligibility, unchanged keyed content fingerprint and suppression state while the twelve-hour snapshot is still active.
For Mastodon, Hacker News and YouTube, that selection returns current public API text, display attribution, source name, broad publication time and a validated canonical source URL. HEARD renders the text as escaped text—never provider HTML, scripts or an iframe—so the voice can be read without leaving the field. OPEN ORIGINAL is a separate optional visit. The resolved response is no-store and is not persisted or logged. Bluesky retains its provider-owned official embed boundary after the same explicit selection, with a validated direct-link fallback.
A field, not an archive
Raw API material and candidate text remain only in one bounded refresh or voice-resolution operation, for at most thirty seconds; text is not placed in PostgreSQL, application logs, public signals or snapshots. OpenAI receives only a bounded unseen editorial batch with store:false, but provider abuse/security retention remains governed by the operator’s OpenAI data-control settings and cannot be described as guaranteed zero retention. HEARD stores a bounded seven-day reservoir containing provider, opaque encrypted source reference, keyed content/account/reference fingerprints, fixed decision values and timestamps—without raw text, display attribution or source URL. Decisions from older rubric versions are retained only until their own expiry and are never reused. Eligible voices come from the previous seven days and must remain age-eligible through the snapshot’s hard expiry plus five minutes. A snapshot is fresh for six hours, remains available but stale from six to twelve hours, and becomes unavailable at twelve hours. A successful refresh atomically replaces the previous snapshot only when it passes the quality contract. Failure or insufficient quality preserves a still-valid prior snapshot without extending its deadline. Expired reads fail closed even if cleanup is delayed.
Limits and failures
A normal field read reads only the authoritative PostgreSQL snapshot and does not start collection or model work. A bodyless authenticated operator endpoint obtains one distributed lease before a bounded refresh. If a snapshot is absent or expired, HEARD shows unavailable. Test fixtures exist only in test-owned and local controlled paths and are excluded from deployment artifacts. Public API responses use no-store so a CDN cannot prolong a removed field.
Model processing is capped at thirty unseen candidates, one finalized call per deterministic six-hour UTC bucket and eight calls per UTC day, with database-enforced request, input-token, output-token and cost reservations. A bounded seven-day reservoir stores only opaque encrypted source references, keyed fingerprints and fixed editorial decisions—never text, display attribution or source URLs. When cadence or budget is exhausted, HEARD makes no model request, ignores unseen candidates and may publish revalidated quality source expressions; a field is replaced only when the full quality contract is met. That hard contract is three voices from three provider-scoped accounts, at least one strong voice and no more than two considered opinions. Provider and topic diversity beyond that are preferences; a valid single-provider field is labelled narrow. Otherwise a still-valid field remains until its own expiry. Selection remains fallible and shaped by language, topic rules, queries, hashtags, instance settings, recent-story ranking and API availability. The field is neither a survey nor a public-opinion estimate. Resolving a voice makes one bounded server-to-provider request; only choosing OPEN ORIGINAL creates a browser-to-provider visit. If revalidation discovers that an opened source is no longer eligible, that signal is removed immediately; unrelated voices remain unless the hard field minimum no longer holds. A verified removal remains suppressed by keyed HMAC.