Cosmo Chat: AI-Powered Documentation Assistant
Architected and built an intelligent documentation assistant for Connect Design System, the most deployed design system for internal products at JP Morgan Chase & Co., powering bankers, analysts, and various roles across private bank and asset wealth management.
Cosmo Chat
What I built
An AI-powered documentation assistant for Connect Design System at JP Morgan Chase. Connect is the most deployed design system for internal products, used by bankers, analysts, and product teams across private bank and asset wealth management.
Role: Design Systems Engineer & Technical Lead Stack: React, Python Flask, MCP Protocol, Nielsen's Heuristics Engine
Why it needed to exist
Our documentation site was a mess. Developers couldn't find what they needed. The search was bad, the information architecture was worse, and the UX of actually navigating docs made people give up. So they'd come to us instead. Teams messages, office hours, Zoom calls, in-person walk-ups, all asking the same questions: "What does this component do?" "How do I use this prop?" "What's the right pattern for this layout?"
I kept answering the same questions. The information existed, it was just impossible to find. So I started thinking about what it would look like to have a chat interface that could answer these questions instantly, and actually help people ideate on patterns and understand what they could build with our components.
How it started
Cosmo started as a 45-line Python script. It read from a JSON file I'd manually filled with common questions (categorized by code, documentation, patterns, UX, office hour questions, and misc) and returned answers. That was it. No AI, no semantic search, no MCP. Just a script and a JSON file.
But it worked well enough that people started using it, and every time it couldn't answer something, I'd add to the JSON. Over time, the gaps in the JSON file became the roadmap for what to build next.
What I built on top of it
The fallback system
The main architecture decision was building a cascading fallback system. When a query comes in, Cosmo tries the fastest, most reliable source first and only reaches for slower external sources when needed:
def process_message(message, category=None):
# Priority 1: Exact match in the JSON knowledge base (~50-100ms)
if exact_match_found:
return exact_answer
# Priority 2: Semantic similarity search (~100-200ms)
if semantic_match_found:
return semantic_answer
# Priority 3: External MCP server for design system docs (~300-500ms)
if mcp_response_available:
return mcp_response
# Priority 4: Suggest related questions
return similar_questions_or_guidanceMost queries still hit the original knowledge base. It's fast and accurate because I curated it from real questions. The MCP integration was the big unlock: instead of manually adding every answer, Cosmo could pull directly from the Connect Design System documentation and learn from it. I spent a lot of time giving the MCP server enough context to return useful answers, not just raw doc dumps.
Component-aware routing
The system detects when someone's asking about a UI component and routes to specialized handlers:
component_keywords = ["component", "button", "card", "dropdown"]
prop_keywords = ["props", "properties"]
usage_keywords = ["how to use", "example"]
if has_component and (has_prop or has_usage):
connect_response = try_connect_mcp_fallback(message)Component questions need specific, structured answers. Props, examples, do's and don'ts. Not general paragraphs.
PDF processing
Internal docs lived in PDFs with tables, images, and messy layouts. I built multi-modal extraction that could handle all of it:
def extract_text_from_pdf(pdf_path):
text = page.get_text("text") # Layout-preserving extraction
table_text = extract_tables_from_blocks(blocks) # Table detection
img_text = pytesseract.image_to_string(image) # OCR for imagesDocuments get chunked at semantic boundaries with overlap so answers maintain context even when information spans multiple pages.
Design critique engine
I also built an automated design critique tool based on Nielsen's 10 usability heuristics:
class ContentScorer:
def __init__(self):
self.weights = {
'sentence_case': {'headings': 2.0, 'paragraphs': 1.0, 'buttons': 1.0},
'writing_issues': {'headings': 1.0, 'paragraphs': 2.0}
}
def calculate_score(self, analysis_results, text_content):
# Weighted scoring, returns 0-10 with breakdownThis let developers get instant feedback on their UI implementations without waiting for a design review.
What was hard
Query ambiguity
People ask the same question a dozen different ways. "What are button props?" and "How do I use buttons?" and "Button properties?" all need the same answer. I built multi-tiered matching: exact strings first, then substring detection, then keyword extraction, then semantic similarity. That's what it takes to handle how people actually ask questions.
Speed vs. accuracy
Searching multiple sources sequentially took 3-5 seconds, which felt broken. The fix was the priority ordering: check the fast local JSON first, only fall back to MCP when needed, and cache aggressively. Most queries resolve in under 200ms from the local knowledge base.
Results
The numbers are rough estimates. I didn't have a formal analytics pipeline, but based on usage logs and team feedback:
- Query success rate went from around 60% to well over 90% - Response times under 500ms on average (most under 200ms from local KB) - Cut documentation search time roughly in half - Noticeably fewer repeat questions in Teams and office hours - Implementation errors dropped significantly once prop correction was working
The system handles thousands of queries daily and became the default way developers interact with Connect documentation.
What I'd take away from this
It started as 45 lines of Python reading from a JSON file. Every feature got added because a real gap showed up, not because I planned it upfront. The JSON file's missing answers became the spec for semantic search. Semantic search not being enough became the spec for MCP integration. MCP's raw output became the spec for component-aware routing.
Start with the simplest thing that works, watch where it breaks, build the next layer when you actually understand the problem.
Architecture overview
User Query
|
Frontend (React) - Natural language interface with syntax highlighting
|
API Gateway (Flask) - Routing and orchestration
|
|-- Knowledge Base (JSON, curated from real questions)
|-- PDF Search (Multi-modal extraction + semantic chunking)
|-- MCP Server (Connect Design System documentation)
|-- Design Critique (Nielsen's heuristics engine)
|
Response - Formatted with code examples and citationsStack rationale: React frontend for component reusability and integration with the existing design system. Python Flask for fast prototyping and access to NLP/document processing libraries. MCP Protocol for standardized connection to external documentation sources.