schedule a call
← All posts

Automating Competitive Keyword Reports: From Scrape to Slack in One Pipeline

October 8, 2026by Marco CoronadoArtificial Intelligence
A diagram showing a data pipeline flowing from app store scraping through keyword comparison to a Slack notification

Manual competitive keyword research is a tax you pay every week. Someone on your team pulls up the App Store, screenshots competitor listings, pastes text into a spreadsheet, and tries to remember what changed since last Tuesday. That process breaks the moment the person doing it gets pulled onto something else — and the insight it produces is already stale by the time it lands in a Slack message.

AI automation eliminates that entire loop. This guide walks through a concrete pipeline: scrape competitor app metadata on a schedule, diff the keyword signals against last week's snapshot, and deliver a structured digest directly to Slack. No dashboards to log into, no manual work, no forgotten Monday reports.

What the Pipeline Actually Does

Before writing a line of code, it helps to be precise about scope. This pipeline does four things:

  1. Fetches public App Store and Google Play metadata for a defined list of competitor app IDs — title, subtitle, short description, long description, keywords field (iOS only, when visible via certain endpoints), and developer name.
  2. Compares the current snapshot against the previous week's snapshot, flagging new terms, removed terms, and title/subtitle changes.
  3. Runs the diff through a language model to summarize what the changes likely signal — a seasonal push, a new feature, a repositioning — in plain English.
  4. Posts a formatted digest to a Slack channel with grouped findings and a link to the full diff stored in cloud storage.

That's it. The pipeline doesn't do rank tracking (use a dedicated ASO tool for that), it doesn't automate your response to the findings, and it doesn't replace human judgment about what to do next. It just makes sure the raw intelligence arrives automatically, every week, whether or not anyone remembered to look.

Pipeline Architecture at a Glance

Stage Tool / Service What it produces
Scheduler GitHub Actions (cron) or AWS EventBridge Triggers the run weekly
Scraper Python + requests / app-store-scraper lib Raw JSON metadata per app
Storage AWS S3 (semnexus-blog-media bucket pattern) Timestamped snapshots
Diff engine Python deepdiff or custom set comparison Keyword delta object
LLM summarizer OpenAI gpt-4o-mini via API Human-readable summary
Notifier Slack Incoming Webhook Weekly digest message

The whole thing runs in a single Python script invoked by the scheduler. No persistent server, no database, no web UI. If you want to harden it into a proper service later, that's a separate decision — but starting serverless keeps operational overhead near zero.

Step 1: Scrape the Metadata

The app-store-scraper Python library wraps the iTunes Search API and the Google Play internal API. It's not official and it will occasionally break when Apple or Google changes their response format — plan for that by wrapping every call in a try/except and logging failures rather than crashing the run.

from app_store_scraper import AppStore, PlayStore

def fetch_ios(app_id: str, country: str = "us") -> dict:
    app = AppStore(country=country, app_name="", app_id=app_id)
    # The library returns listing metadata without reviews when you call .review()
    # For metadata only, hit the lookup endpoint directly
    import requests
    r = requests.get(
        f"https://itunes.apple.com/lookup?id={app_id}&country={country}",
        timeout=10
    )
    r.raise_for_status()
    results = r.json().get("results", [])
    return results[0] if results else {}

Pull the fields you care about: trackName, subtitle (iOS 11+), description, releaseNotes, primaryGenreName, averageUserRating. Store the raw response as a timestamped JSON file in S3 — {app_id}/{YYYY-WW}.json where WW is the ISO week number. Keeping the full raw response means you can re-derive any field later without re-scraping.

For Google Play, the play_scraper library gives you title, description, short_description, score, installs, and recent_changes.

Don't try to scrape the iOS keywords field. Apple doesn't expose it through any public endpoint. What you can do — and what actually matters more — is extract the keyword signals from the title, subtitle, and description text itself.

Step 2: Extract Keyword Signals from Text

Raw description text isn't a keyword list. You need to normalize it into a comparable set of terms before you can diff it week over week.

A simple approach that works in practice:

import re
from collections import Counter

STOPWORDS = {"the", "a", "an", "and", "or", "for", "in", "of", "to", "with", "your", "you"}

def extract_terms(text: str, top_n: int = 60) -> set:
    tokens = re.findall(r'\b[a-z]{3,}\b', text.lower())
    filtered = [t for t in tokens if t not in STOPWORDS]
    counts = Counter(filtered)
    return {term for term, _ in counts.most_common(top_n)}

Run this against the concatenation of title + subtitle + short description for each competitor. The title and subtitle carry disproportionate ASO weight, so you can optionally double-weight them by repeating those tokens before the concatenation.

Store both the raw JSON and the extracted term set. The term set is what you diff; the raw JSON is your audit trail.

Step 3: Diff Last Week Against This Week

Load both snapshots from S3, compute the set difference, and structure the result.

import boto3, json

s3 = boto3.client("s3")

def load_snapshot(bucket: str, app_id: str, week: str) -> set:
    key = f"{app_id}/{week}_terms.json"
    try:
        obj = s3.get_object(Bucket=bucket, Key=key)
        return set(json.loads(obj["Body"].read()))
    except s3.exceptions.NoSuchKey:
        return set()

def compute_diff(current: set, previous: set) -> dict:
    return {
        "added": sorted(current - previous),
        "removed": sorted(previous - current),
        "retained": len(current & previous),
    }

Also check for title or subtitle changes directly — those are high-signal moves worth calling out explicitly regardless of the term-level diff.

def metadata_changes(current_raw: dict, previous_raw: dict) -> list:
    changes = []
    for field in ["trackName", "subtitle", "shortDescription"]:
        c, p = current_raw.get(field, ""), previous_raw.get(field, "")
        if c != p:
            changes.append({"field": field, "from": p, "to": c})
    return changes

Step 4: Summarize with an LLM

A list of added and removed terms is useful for an ASO analyst but not for most of the people on your Slack channel. Pass the diff to a language model and ask it to interpret the signal.

import openai

def summarize_diff(app_name: str, diff: dict, metadata_changes: list) -> str:
    prompt = f"""
You are an ASO analyst. A competitor app called "{app_name}" made the following changes to its App Store listing this week.

Added keywords: {diff['added']}
Removed keywords: {diff['removed']}
Metadata field changes: {metadata_changes}

In 3-4 sentences, explain what these changes likely signal about the competitor's current marketing focus or product strategy. Be specific. Do not use filler phrases.
"""
    response = openai.chat.completions.create(
        model="gpt-4o-mini",
        messages=[{"role": "user", "content": prompt}],
        max_tokens=200,
        temperature=0.3,
    )
    return response.choices[0].message.content.strip()

Keep temperature low — you want analysis, not creativity. gpt-4o-mini is sufficient here and keeps costs minimal; at roughly 400 tokens per competitor per week, you're spending pennies even across a list of 20 apps.

Want the scraping, diffing, and summarization handled for you? Semnexus builds end-to-end AI automation pipelines for app marketing teams. See what we put together at /mobile-app-marketing-services.

Step 5: Post to Slack

Use a Slack Incoming Webhook. Format the message with Block Kit so it's readable rather than a wall of text.

import requests as http_requests

def post_to_slack(webhook_url: str, summaries: list[dict]) -> None:
    blocks = [
        {"type": "header", "text": {"type": "plain_text", "text": "📊 Weekly Competitor ASO Digest"}},
        {"type": "divider"},
    ]
    for s in summaries:
        blocks.append({
            "type": "section",
            "text": {
                "type": "mrkdwn",
                "text": f"*{s['app_name']}*\n{s['summary']}\n_Added: {', '.join(s['added'][:5]) or 'none'} | Removed: {', '.join(s['removed'][:5]) or 'none'}_"
            }
        })
    http_requests.post(webhook_url, json={"blocks": blocks}, timeout=10)

Cap the inline term lists at five items — the Slack message is a trigger to investigate, not the full report. Link to the S3-stored JSON diff for anyone who wants to dig deeper.

Failure Modes to Watch

This pipeline is simpler than a multi-step AI agent, but it still fails in predictable ways. The most common issues in our engagements:

  • Scraper drift: App Store and Play Store response shapes change without notice. Pin your scraper library version and add a schema validation step that alerts you if expected fields go missing.
  • Empty diffs that aren't empty: If a competitor updates their description without changing the top-60 terms, the diff will show nothing. Tune top_n or add a raw character-count change check as a fallback signal.
  • LLM hallucinations on thin data: If added and removed are both empty but there's a title change, the model can overinterpret a minor edit. Add a guard: only call the LLM if the diff has at least three changed terms or a metadata field change.
  • Rate limits and blocks: Scrape politely. Add a 1–2 second sleep between requests. Run at an off-peak hour (the GitHub Actions cron runs at 12:00 UTC daily for Semnexus's own scheduled posts — a Sunday night scrape fits naturally).

For a deeper look at how production AI pipelines break in practice, the post on agent failure modes covers the failure taxonomy in detail — most of it applies here even though this pipeline isn't technically an agent.

FAQ

How many competitor apps can I track before costs become a problem?

At gpt-4o-mini pricing, tracking 50 apps costs approximately $0.50–$1.00 per week in LLM calls. Scraping is free (public endpoints). S3 storage for a year of JSON snapshots across 50 apps is negligible. The real constraint is scraper reliability at scale, not cost.

Can I track Google Play and App Store in the same run?

Yes. Structure your config as a list of objects with an app_id and platform field. The scraper functions are separate; the diff and summarization logic is identical.

What if a competitor hasn't changed anything?

Skip the LLM call and post a one-liner: "No significant changes detected." Don't waste tokens summarizing nothing, and don't clutter the Slack digest with empty entries.

How is this different from a paid ASO tool?

Paid ASO tools (AppFollow, AppTweak, Sensor Tower) give you rank tracking, review analytics, download estimates, and historical data you can't reconstruct yourself. This pipeline gives you a lightweight, customizable, automated diff that you fully control — no per-seat pricing, no vendor lock-in, and you can extend it however you want. The two are complementary, not competing.

Can I add email delivery instead of Slack?

Replace the Slack webhook call with an SMTP send or a SendGrid API call. The rest of the pipeline is identical. You can also run both in parallel.

How do I handle apps that get delisted or change their App Store ID?

Add a validation step after the scrape: if the API returns zero results for a known app ID, flag it explicitly in the Slack message rather than silently dropping that competitor from the diff. Relisted apps sometimes get new IDs, so keep a mapping file you can update manually.


If you want this pipeline running without building it yourself — or you want it extended into a full competitive intelligence system with rank tracking, review sentiment, and automated response recommendations — talk to the Semnexus team. Or skip straight to booking a 30-minute call and we'll scope it out together.

lets connect

SEM Nexus is ready to help you find unique solutions for your app. Get in touch to learn more about your project and receive the full SEM Nexus treatment.

By partnering with SEM Nexus, you can confidently launch your app and get your product into the hands of customers, achieving unparalleled mobile growth.

get in touch now!
breaker
logo 98 Cuttermill Road STE 223N,
Great Neck, New York, 11024
follow us
facebookinstagramlinkedin
our newsletter
subscribe!