Marketplace / Blueprint 09
LIVEAgents / agency

Research Intelligence

Search, crawl, extract, structure and retain web intelligence as reusable organizational knowledge.

Review my build

Interactive platform map

Architecture in context

The focused blueprint, its required foundation, and declared recommendations.

Core Selected Automatic Required path
Blueprint 09 / Agents

Research Intelligence

Search, crawl, extract, structure and retain web intelligence as reusable organizational knowledge.

2 vCPU3 GB RAM5 services
01 / Problem

What this replaces

Web research is repetitive, disappears into individual sessions and often relies on expensive or opaque search APIs.

02 / Outcome

What your team gains

A private research pipeline that discovers sources, crawls approved targets, structures findings and stores them for semantic/graph retrieval.

03 / Capability

What is inside the blueprint

Private metasearch across many engines/categories
HTTP search API with JSON/CSV/RSS output options
Reduced user profiling and query leakage
Concurrent/fault-tolerant Scrapy crawling
CSS/XPath/regex extraction
Spiders and generic sitemap/feed spiders
Scheduler, downloader and middleware extensibility
Cookies, auth, caching, compression, robots and crawl depth features
AutoThrottle/concurrency controls
Item pipelines for cleaning/validation/persistence
JSON/JSONL/CSV/XML feed exports and multiple storage backends
Container-per-crawl isolation in AICORTEX
Embedding and graph enrichment after crawl
04 / Architecture

How it fits the platform

Agent -> SearXNG discovery -> scrapy-mcp isolated crawl -> extraction/pipelines -> Postgres/pgvector + Neo4j -> retrieval/analysis.

Included services

SearXNG, scrapy-mcp, Scrapy, PostgreSQL/pgvector, Neo4j

Platform requirements
05 / Delivery

From prerequisites to operation

Prerequisites
  1. Research scope
  2. Target-domain policy/terms review
  3. Storage schema
  4. Crawler templates
Deployment
  1. Configure SearXNG API
  2. Deploy/rebuild pinned Scrapy MCP image
  3. Create crawl isolation workflow
  4. Define extraction pipelines
  5. Connect persistence/indexing
Configuration
  1. Search engines/categories
  2. Crawl depth
  3. Concurrency and delay
  4. Robots compliance
  5. Extraction selectors
  6. Deduplication
  7. Export/storage targets
Operations
  1. Target-site changes
  2. Rate limits/bans
  3. SDK/image dependency drift
  4. Crawler resource isolation
  5. Data freshness
06 / Combinations

What this unlocks with other layers

Research Intelligence + Knowledge Hub & RAG + Knowledge Graph & Institutional Intelligence

Research-to-Knowledge

Discovery becomes a retained, structured research corpus instead of disappearing after one chat.

07 / Technology

Technology behind this capability

SearXNGRUNNINGpgvectorRUNNING in enterprise DB imageNeo4jRUNNING - 5-community in censusModel Context Protocol (MCP)MULTIPLE RUNNING SERVERSScrapyRUNNING through scrapy-mcpscrapy-mcpAICORTEX native