Automation & Digital Systems

Multi-Source AI Research Pipeline

Designed and implemented a scheduled research pipeline that collects AI developments from RSS and API sources, normalizes them into a shared schema, filters and deduplicates records, uses Gemini for structured relevance analysis, and stores useful findings in a searchable Notion research library.

Role
Workflow Designer & AI Automation Developer
Status
Implemented · Production workflow published
Environment
Local n8n via Docker + Cloud APIs
Sources
TechCrunch AI RSS · Product Hunt GraphQL API
AI
Google Gemini
Destination
Notion
n8nAI AutomationGeminiGraphQLNotionDocker

Converted repetitive AI research into a scheduled, stateful pipeline that continuously discovers, evaluates, deduplicates, and organizes useful developments.

Project Overview

Keeping up with AI developments means monitoring several sources, separating useful updates from noise, avoiding repeated content, and organizing findings for later reference. Doing that manually quickly becomes repetitive.

I built an n8n pipeline that collects new items, converts different source formats into one structure, prioritizes what should be evaluated, uses Gemini for structured analysis, and stores relevant findings in Notion.

The goal was a small reliable information system: aware of completed work, conservative with API quota, resistant to duplicates, and capable of retrying after external-service failures.

Implementation scope

RSS ingestionGraphQL integrationSource normalizationStructured Gemini outputPersistent stateNotion integrationRetry behaviorProduction publishing

The Manual Problem

A typical research session required checking multiple sources, opening individual posts, judging AI relevance, writing summaries, and transferring notes into a database.

The automation therefore needed to solve more than aggregation. It had to decide what deserved AI processing, what had already completed, and what should happen after partial failure.

01

Different data formats

RSS and GraphQL expose different field structures.

02

Repeated content

Previously reviewed URLs can appear in later runs.

03

Limited AI quota

Not every collected item deserves an LLM request.

04

Duplicate records

Repeated writes reduce trust in the research library.

05

Failure recovery

Incomplete work must remain eligible for a later retry.

Design priorities

Multi-source

One downstream workflow for fundamentally different sources.

Quota-aware

Filter, deduplicate, and prioritize before Gemini.

Stateful

Remember only URLs that reached a completed outcome.

Duplicate-safe

Protect both AI capacity and the final Notion database.

Retry-ready

Leave incomplete work eligible for a future run.

Solution Architecture

Source-specific work happens at the edge: TechCrunch enters through RSS, while Product Hunt is retrieved through GraphQL and pre-filtered for AI signals.

After both sources are normalized and merged, every item follows the same state check, prioritization, AI analysis, duplicate protection, persistence, and terminal logging path.

Multi-source AI research pipeline architecture

Scheduled Trigger

Starts a controlled research run

Source-specific ingestion

TechCrunch AI

RSS feed

RSS Transformation

Maps feed fields to the shared schema

Product Hunt

Authenticated GraphQL API

AI Pre-Filter

Removes obviously unrelated products

API Transformation

Maps post fields to the shared schema

Data Normalization

One source-neutral research shape

Merge Research Sources

One downstream item stream

Persistent Processing Check

Removes URLs with completed outcomes

Sort by Publication Date

Newest eligible items first

Daily Processing Limit

Maximum five items per run

Gemini AI Analysis

Structured relevance and research context

Structured Research Record

Metadata and AI output combined

Relevant Research?

Routes the terminal outcome

No · Irrelevant

Processing Log

Records the irrelevant terminal outcome

Yes · Relevant

Notion Duplicate Check

Protects the destination independently

Existing

Processing Log

Records already_in_notion

New

Save to Notion

Creates the research page

Processing Log

Records saved_to_notion

Source-specific collection is handled at the edge of the workflow. After normalization, every item follows the same processing, AI analysis, duplicate protection, and storage logic.

How the Pipeline Works

Five stages turn mixed source data into consistent, searchable research.

  1. 01 — COLLECT

    A scheduled n8n run retrieves TechCrunch AI through RSS and Product Hunt posts through authenticated GraphQL requests.

  2. 02 — PRE-FILTER

    A lightweight keyword check removes obviously unrelated Product Hunt posts before any limited AI quota is used.

    Cheap deterministic filtering happens before expensive AI reasoning.

  3. 03 — NORMALIZE

    Both source formats are mapped to the same source-neutral fields, so downstream nodes do not need source-specific logic.

    sourcesource_typetitleurlpublished_atauthorraw_excerpt
  4. 04 — PRIORITIZE

    Completed URLs are removed, eligible records are sorted newest first, and only the first five proceed to analysis.

  5. 05 — ANALYZE & PERSIST

    Gemini returns structured analysis; relevant, genuinely new items reach Notion, and every completed path records a terminal outcome.

System in Action

The native diagram explains the architecture; these screenshots show the published workflow and the data it produces.

Full Workflow

Full size
Full n8n production workflow for the Multi-Source AI Research Pipeline
The production workflow orchestrates source ingestion, normalization, AI analysis, duplicate checks, and Notion persistence inside a single scheduled n8n pipeline.

Source Ingestion Detail

Full size
n8n ingestion workflow showing TechCrunch RSS and Product Hunt API normalization before merging
The ingestion layer merges TechCrunch AI RSS content with Product Hunt API results after source-specific filtering and normalization.

AI Research Library

Full size
Notion AI Research Library containing structured findings from TechCrunch AI and Product Hunt
Relevant items are written to Notion with category, importance, summary, practical significance, intended audience, source, and publication date.

Processing Log

Full size
n8n processing log with irrelevant, already in Notion, and saved to Notion terminal statuses
A separate log captures terminal outcomes such as irrelevant, already_in_notion, and saved_to_notion, providing persistent state and safe retries.

Quota-Aware AI Processing

Development used a relatively small Gemini free-tier request allowance, so quota became part of the processing design rather than an afterthought.

Completed records are removed first. The remaining items are sorted newest to oldest, capped at five, and processed individually with a delay and automatic retries for temporary failures.

Designed sequence

Deduplicate
Prioritize by freshness
Limit to five
Gemini analysis

Wasteful sequence avoided

Gemini analysis
Deduplicate later

AI Analysis & Structured Output

Predictable AI response

Gemini evaluates each eligible item and returns a constrained structure rather than unrestricted prose. Later nodes can route and store the response without interpreting free-form text.

{

"relevant": boolean,

"summary": string,

"why_it_matters": string,

"who_its_for": string,

"category": enum,

"importance": 1–10,

}

Allowed categories

AI ModelsAI Agents & AutomationAI CodingAI ProductivityAI MediaAI ResearchAI Business & IndustryOther
Canonical research record

A transformation node merges the source metadata with Gemini’s response. This source-neutral record is the contract used by both Notion and the processing log.

sourcesource_typetitleurlpublished_atauthorrelevantsummarywhy_it_matterswho_its_forcategoryimportance

State, Retry & Duplicate Protection

The workflow distinguishes between an item it has seen and one it has successfully finished processing. A URL enters the persistent processing log only after a terminal outcome.

Completed outcomes

irrelevantalready_in_notionsaved_to_notion

Irrelevant items are still completed work. Relevant items receive a second URL check inside Notion, so the final library stays clean even if processing history and destination state ever diverge.

Completed

New URL
Gemini / Notion
Terminal outcome
Write processing log

Incomplete

New URL
Gemini / Notion failure
No terminal log
Eligible for retry

Layer 1 — Processing Log

Before Gemini: “Has this URL already completed?”

Avoids repeated AI processing across scheduled runs.

Layer 2 — Notion

Before storage: “Does this URL already exist?”

Protects destination integrity independently.

“Seen” and “successfully finished” are different workflow states. Preserving that distinction keeps failed work recoverable.

Metrics & Operational Value

2
Live research sources
2
Duplicate-protection layers
3
Tracked terminal outcomes
5
Maximum AI evaluations per run

Automated daily intake

A scheduled workflow collects new research without repeated manual source checking.

Multi-source ingestion

RSS and authenticated API data enter the same normalized pipeline.

Focused AI usage

Deterministic filtering and freshness ranking reserve Gemini for eligible work.

Consistent output

Every accepted item reaches Notion with the same structured research fields.

Duplicate-safe persistence

Processing history and a destination lookup protect two different failure points.

Operational reliability

Terminal logging supports repeated schedules and safe retries after incomplete runs.

Reflection & Lessons Learned

Connecting feeds and APIs to an LLM was straightforward. The more valuable work was deciding when an item becomes complete, how different formats become interchangeable, where duplicate protection belongs, and what should happen when processing stops halfway through.

  • Treat API limits as an architectural constraint, not a late optimization.
  • Normalize different sources before shared analysis and persistence.
  • Track completed work rather than every item the workflow has merely seen.
  • Keep an operational log separate from the human-facing destination database.
  • Maintain development and rollback versions so experiments stay out of the scheduled production workflow.

Those decisions transformed a chain of integrations into a small, resilient information pipeline.

Back to Projects
© 2026 Ken Gilmer P. Macawili. All Rights Reserved.
0%