# Build an X Crypto Narrative Velocity Scanner

Monitor curated crypto accounts, detect accelerating narratives, and send only the strongest candidates to Claude for interpretation.

## What you will achieve

You will have a Python scanner that stores X posts in SQLite, ranks token narratives by velocity and independent convergence, flags possible campaigns, and produces a compact JSON shortlist for review.

## Who this is for

Crypto researchers, traders, and technical analysts who are comfortable running basic terminal commands and editing JSON.

**Difficulty:** Intermediate

## Short tutorial

Monitor curated crypto accounts, detect accelerating narratives, and send only the strongest candidates to Claude for interpretation.

### Guide: Build an X Narrative Velocity Scanner

**Time:** 45 minutes for setup, followed by a 30 d

## What you will build

This scanner automates research across a curated set of X accounts. Instead of ranking narratives by total mentions, it looks for entities whose discussion is accelerating relative to their own historical baseline.

The workflow is:

1. Ingest recent posts from selected accounts.
2. Extract tickers and configured project aliases.
3. Store posts and mentions in persistent SQLite tables.
4. Score velocity, independent account convergence, and novelty.
5. Flag cold starts and suspiciously similar phrasing.
6. Send a shortlist of up to 20 candidates to Claude for interpretation.

The supplied implementation uses `twitterapi.io`, not the official X API. Its vendor specific logic is isolated in `fetch_user_tweets()`, so you can replace that function if you choose another data source. Third party data providers may create reliability or terms of service risks, so evaluate the provider before using it.

## 1. Prepare the project

Use the runnable files from the `alpha-scanner` archive:

```text
alpha-scanner/
├── config.json
├── README.md
└── scanner.py
```

Do not copy the code from the Word version. Its indentation, comments, and `__name__` entry point were damaged by document formatting.

From the project directory, install the only external Python dependency:

```bash
pip install requests
```

Set the API key and choose a durable database location:

```bash
export TWITTERAPI_KEY=your_key
export SCANNER_DB=/persistent/path/alpha.db
export SCANNER_CONFIG=/path/to/alpha-scanner/config.json
```

The database must persist between runs. The 30 day baseline, cursors, extracted mentions, and alert history all live in SQLite. An ephemeral environment that starts with a fresh database cannot calculate useful velocity.

## 2. Curate the accounts

Open `config.json` and edit the `accounts` array. Handles must not include the `@` symbol.

Start with roughly 10 to 20 accounts rather than importing all 103 supplied accounts immediately. A smaller set makes it easier to inspect data quality, control API usage, and identify correlated groups.

Here is the supplied seed list:

```json
"accounts": [
  "milesdeutscher",
  "cobie",
  "hsakatrades",
  "0xngmi",
  "smyyguy",
  "CryptoKaleo",
  "AlgodTrading",
  "0xdoug",
  "TheFlowHorse",
  "Pentosh1"
]
```

The curated source list contains 103 possible contributors. Examples you can substitute or add include:

```json
"accounts": [
  "0xdoug",
  "TheFlowHorse",
  "Pentosh1",
  "DefiSquared",
  "stacy_muur",
  "0xJeff",
  "alpha_pls",
  "ViktorDefi",
  "gametheorizing",
  "krugermacro",
  "KobeissiLetter",
  "fejau_inc"
]
```

Treat any initial list as a hypothesis. Account quality should eventually be measured by whether an account posts before a move begins, not by follower count or whether prices rose sometime after a post.

## 3. Configure entity extraction

The scanner recognizes explicit cash tags such as `$HYPE`. It also maps project names and aliases to canonical identifiers through the `entities` object.

Use this copy-ready starting configuration:

```json
{
  "accounts": [
    "0xdoug",
    "TheFlowHorse",
    "Pentosh1",
    "DefiSquared",
    "stacy_muur",
    "0xJeff",
    "alpha_pls",
    "ViktorDefi",
    "gametheorizing",
    "krugermacro"
  ],
  "window_hours": 6,
  "baseline_days": 30,
  "novelty_days": 30,
  "min_mentions": 3,
  "min_z": 2.0,
  "min_convergence": 2.0,
  "shortlist_size": 20,
  "max_pages_per_account": 3,
  "sleep_between_accounts": 0.2,
  "phrase_overlap_flag": 0.35,
  "onset_threshold": 2.0,
  "ticker_stoplist": [
    "usd",
    "usdt",
    "usdc",
    "btc",
    "eth",
    "sol",
    "the",
    "a",
    "it"
  ],
  "entities": {
    "HYPE": ["hyperliquid", "hype"],
    "ARROW": ["arrow protocol", "arrow"],
    "JUP": ["jupiter", "jup"],
    "TIA": ["celestia", "tia"],
    "ENA": ["ethena", "ena"]
  },
  "account_weights": {}
}
```

### How the main settings work

| Setting | Purpose |
|---|---|
| `window_hours` | Defines the current activity window. The supplied default is six hours. |
| `baseline_days` | Sets the historical comparison period. |
| `novelty_days` | Suppresses entities that already generated a recent alert. |
| `min_mentions` | Requires at least this many mentions in the current window. |
| `min_z` | Requires activity to exceed the entity's normal baseline. |
| `min_convergence` | Requires support from sufficiently independent accounts. |
| `shortlist_size` | Limits the JSON output sent to the model. |
| `phrase_overlap_flag` | Controls when similar language is flagged as a possible campaign. |
| `ticker_stoplist` | Excludes majors and ambiguous words that create noise. |
| `entities` | Maps names and aliases to one canonical entity. |

Expand the alias table as you observe missed or inconsistent matches. Short aliases under four characters are not matched as ordinary words by the supplied extractor, although cash tags can still be detected.

## 4. Start ingestion

Run the first ingestion manually:

```bash
python scanner.py ingest
```

A successful run prints output similar to:

```text
ingested 180 new posts
extracted 24 mentions from 180 posts
```

The scanner stores:

* Raw posts, including handle, timestamp, text, likes, and URL.
* Canonical entity mentions.
* Per-account cursors, which prevent repeated ingestion.
* Alerts and their full payloads.

Schedule ingestion every 15 minutes. For example, a cron entry could be:

```cron
*/15 * * * * cd /path/to/alpha-scanner && /usr/bin/python3 scanner.py ingest >> scanner.log 2>&1
```

Ensure the scheduled process receives `TWITTERAPI_KEY`, `SCANNER_DB`, and `SCANNER_CONFIG`. Depending on the host, that may require defining them in the cron environment or loading them from a protected shell script.

## 5. Generate the shortlist

Run scoring after ingestion:

```bash
python scanner.py score
```

The output is JSON. Each candidate can include:

* Mentions in the current window.
* Historical baseline mean.
* Velocity z-score.
* Correlation weighted convergence.
* Distinct accounts.
* Account weight.
* Composite score.
* Campaign warning flags.
* Up to three sample posts.

The composite score is driven by velocity, account weight, and convergence. Accounts that historically discuss the same entities at nearly the same time are discounted, because five correlated accounts may represent one conversation rather than five independent confirmations.

The novelty gate suppresses an entity after it has fired within the configured period. This prevents repeated alerts about the same continuing narrative.

## 6. Interpret candidates with Claude

The model should be the final interpretation layer, not the primary detection engine. Sending thousands of raw posts directly to a model is costly and encourages weak pattern finding. Send only the deterministic shortlist.

If the Claude command line tool is configured, use:

```bash
python scanner.py score | claude -p "Here are today's narrative candidates from my tracked accounts. For each candidate, explain what it is, why discussion may be accelerating now, and whether the supplied flags suggest a coordinated campaign. Compare the sample posts, distinguish independent evidence from repeated claims, and rank the candidates by what appears genuinely early. Do not treat mention velocity as proof of investment quality. Return a concise table followed by the key uncertainties for each candidate."
```

For a chat interface, run `python scanner.py score`, paste the JSON, and use this prompt:

```text
You are reviewing a deterministic shortlist from an X narrative velocity scanner.

For each candidate:
1. Explain what the entity or narrative is, using only the supplied evidence.
2. Identify why attention appears to be accelerating now.
3. Compare the distinct accounts and sample posts.
4. Evaluate the cold_start and similar_phrasing flags when present.
5. Separate independent observations from repeated or coordinated claims.
6. State what cannot be verified from the supplied posts.
7. Rank candidates by how genuinely early they appear, not by total popularity.

Do not interpret velocity as proof of quality or future return. Do not invent catalysts, prices, partnerships, or project details absent from the input.

Return columns for Rank, Entity, Evidence of Acceleration, Independence, Campaign Risk, Key Uncertainty, and Research Priority.

SHORTLIST JSON:
[paste scanner output here]
```

A campaign flag is context, not an automatic rejection. The supplied synthetic test correctly surfaced a coordinated looking candidate, but it also produced a false positive when legitimate posts shared a sentence stem. Adjust `phrase_overlap_flag` only after reviewing real results.

## 7. Let the baseline mature

Do not trust early velocity scores. The scanner needs approximately 30 days of persistent history to learn what normal discussion looks like for each entity.

During this period:

* Keep ingestion running consistently.
* Inspect API errors in `scanner.log`.
* Add aliases for missed project references.
* Add ambiguous symbols to `ticker_stoplist`.
* Review whether correlated accounts are being discounted sensibly.
* Preserve every alert so misses remain visible alongside hits.

In the supplied synthetic test, steady JUP and TIA discussion did not surface even though they had more total mentions. ARROW surfaced because it accelerated from near zero across five relatively independent accounts. This is the intended behavior.

## 8. Backtest account quality correctly

The command is:

```bash
python scanner.py backtest
```

However, it will not produce meaningful rankings until you implement the `price_series(entity)` function in `scanner.py`. The supplied function deliberately returns an empty list:

```python
def price_series(entity):
    """PLUG IN YOUR PRICE SOURCE. Return [(datetime, close), ...] or []."""
    return []
```

Your implementation must return ordered `(datetime, close)` pairs. The backtest then compares each account's first mention with the detected onset of a move.

Use lead time rather than simple forward return. An account posting after an asset has already moved should receive no early discovery credit, even if the asset rises further. The supplied scoring also penalizes accounts that mention many entities indiscriminately.

After the backtest works, use its results to populate `account_weights`:

```json
"account_weights": {
  "early_account": 1.4,
  "average_account": 1.0,
  "late_commentator": 0.6
}
```

The source does not prescribe exact weight conversion rules, so introduce weights conservatively and document how you derived them.

## Operational checklist

Before relying on the output, confirm:

* [ ] The runnable archive version of `scanner.py` is installed.
* [ ] `requests` is installed.
* [ ] The API key is available to manual and scheduled runs.
* [ ] `SCANNER_DB` points to durable storage.
* [ ] Account handles omit the `@` symbol.
* [ ] The alias table covers the entities you care about.
* [ ] Ambiguous symbols and unwanted majors are stoplisted.
* [ ] Ingestion runs every 15 minutes without repeated errors.
* [ ] At least 30 days of baseline data has accumulated.
* [ ] Claude receives only the shortlist, flags, and sample posts.
* [ ] Alerts are preserved for later evaluation.
* [ ] `price_series()` is implemented before using backtest results.

The scanner is a research prioritization system, not evidence that an asset is sound or that a trade will be profitable. Its useful output is a smaller, better structured queue of narratives that deserve verification.