Ask Inigo

Personal portfolio


My role
  • Product Design & AI Engineering
Team
  • Solo project
Tools
  • Next.js
  • TypeScript
  • Vercel
  • Vercel AI SDK
  • Gemini 3.5 Flash-Lite
  • Upstash Redis
  • Deterministic tool-based retrieval
Timeline

2026 — Shipped

Description

A grounded AI assistant shipped on this site that answers questions about my projects, experience and skills using only content that already exists here. It refers to me in the third person, retrieves case studies through deterministic tool calls rather than semantic search, and cites the page every answer came from.

Context

Ask Inigo came from the idea of making a portfolio explorable through conversation rather than forcing every visitor to click through every project page. Instead of a generic chatbot, it answers only from the same content module the About and Work pages render from, so what the assistant says and what the site shows cannot drift apart.

Click around...

Challenge

The central challenge was trust.

A portfolio assistant is actively harmful if it confidently invents skills, projects, experience, or achievements that are not actually mine. A recruiter cannot tell an invented claim from a real one, and I am the person who has to answer for it.

The subtler failure is not invention but exaggeration: quietly turning “used Python” into “expert Python developer”, or a university project into production experience. The system had to be able to summarise what my content says without ever strengthening it.

Constraints

It had to run inside free tiers with no payment method configured anywhere, so there is no state in which it can quietly start costing money. When a quota runs out it says so and offers my email rather than falling back to a paid path.

A public LLM endpoint is drainable by anyone with a loop, so the daily allowance needed protecting well enough that the assistant still works when someone who matters visits.

It also had to stay useful when the model is unavailable. With no API key configured the card becomes a static teaser with its input disabled, and the site builds and runs normally.

Research

The obvious approach was retrieval-augmented generation over a vector database — Supabase with pgvector, portfolio content stored as embeddings, semantic search before each answer. I planned that first, then rejected it.

The corpus is about thirteen thousand tokens across twelve projects and three roles, and every one of them is already addressed by a slug. At that size, looking a project up by its slug is exact. Semantic retrieval over the same content can only introduce a chance of returning the wrong project, and it adds a database, a chunking strategy and an embedding refresh step to maintain.

So tool calling became the retrieval layer: the model gets a short briefing listing what exists, and opens the full case study through a function call when it needs one. It is still retrieval-augmented generation, just over a known schema instead of an embedding space.

Deciding not to build the more impressive version was the more useful judgement, and I would rather the portfolio showed that than the machinery.

Iterations

The response renderer started as a visual design I liked, built as a gallery demo. Reading it properly showed it animated a hard-coded string on a fifty-five millisecond timer and looped forever, which in a real thread would have made every finished answer replay itself indefinitely.

Its inline citations had the worse problem: every marker resolved to the first source in the list, so a sentence about one project could display another project’s chip. In a demo with one visible citation that is invisible. On a portfolio it is a false attribution.

I kept the typography, the source chips and the expandable list, and rebuilt the rest. Citations became message-level and are now collected from metadata the tools attach to their own results, so the model is never asked which sources it used and cannot produce a citation for something it did not retrieve.

The streaming is real: text is rendered from the chunks the model actually sends, and nothing waits on a queue. What the original did — reveal a finished string one word every 55ms, then loop — is gone. Gemini turned out to deliver roughly twenty words at a time, so the words carry a bounded entrance instead: each fades in over a sweep capped at 300ms, which is short enough that nobody waits to read and long enough that an answer reads as being written rather than pasted. The cap is the honest part, and raising it would hand the decision about when a visitor may read from the model to the animation.

Key Features

  • Conversational portfolio exploration on the About page
  • Answers grounded in the same content module the site renders from
  • Deterministic project and experience retrieval by slug
  • Real streamed responses, never simulated or paced
  • Citations derived from tool metadata, never written by the model
  • Follow-up questions derived from returned data, so every suggestion is answerable
  • Rate limiting on the platform-reported address, failing closed
  • Third-person identity, refusal preferred to inference
  • One calm message and an email link on every failure path

Final Deliverable

A working assistant on the About page that gives recruiters, collaborators and visitors a faster way to explore my work, built so that the answers stay tied to content I have actually written.

It ships behind two layers of testing: 146 unit tests covering the retrieval, the limiter, the streaming and the renderer, and 45 grounding evaluations run against the live model across hallucination traps, leading questions, prompt injection, refusal to grade, ambiguity and off-topic use. All 45 pass. The unit tests only prove the plumbing works — the evaluations are what say anything about the answers.

Takeaway

The project forced me to separate adding AI because it is visually impressive from adding AI because it improves how information is found.

Most of the work was not the model. It was deciding what the model is not allowed to do: not to speak as me, not to write a link, not to upgrade a claim, and to say it does not know rather than guess. Refusing well turned out to be the feature.

Inline, clause-level citation is the obvious next step, and is deliberately left for a later version: doing it honestly means the model emitting markers that React validates against sources actually retrieved, which is a new fabrication surface to defend rather than a formatting change.