Sree Dayanidhi
Writing RSS riteangle
  • All
  • Riteangle 17
  • Decision systems 10
  • Agent architecture 8
  • Operations 8
  • Context engineering 6
  • Guardrails 6
  • Data platform 5
  • Agent evals 4
  • Agentic architecture 3
  • Advertising 2
  • Marketing 2

Data platform — 5 posts

  • Automating my job hunt outreach with an agent that finds, researches, and drafts — and leaves sending to me

    A director/VP-level job search runs on outreach nobody automates safely, because the risky part is the send. This system automates finding, researching, and drafting end to end, and keeps the one irreversible step out of its code entirely. Underneath it, memory is split three ways: a table store for quantities, an embedding index for meaning, a graph for relationships. The question is no longer how to store data for a human to read, but how to organise it so an agent retrieves the right slice before a model call it pays for. The retrieval half gave the more surprising results: indexing each document's own title bought nearly what the vector index did, four employers' near-identical policies could only be separated by file path rather than by any retriever, and a negative result about rank fusion did not survive re-running it on a larger corpus.

    September 3, 2026 44 min read

  • Five dbt detectors sharing one output contract, and the timezone table that breaks all of them twice a year

    Contact-centre rules about when you may call someone carry real penalties, and they change — so the structural question is whether adding next year's regulation means rewriting this year's checks. Five detectors that know nothing about each other, each emitting the same three fields, means a new rule is a new query rather than a schema migration. The interesting part is that the architecture is sound and the whole thing is still wrong for two months a year, because of a hardcoded timezone table nobody thought was interesting.

    September 1, 2026 5 min read

  • When to use a vector database, when a SQL table, and when a graph, and what it costs to pick wrong

    For decades we modelled data so a human could read it, which is why tabular won. The first consumer now is an agent, and an agent's binding constraint is not legibility but token cost. If context were free you would hand a frontier model all 1,100 documents and ask; it is not, so the job becomes retrieving the smallest correct slice. An agent that drafts outreach kept quoting 600 calls a week from a summary someone wrote once, when the tracker said about 1,850. The fix was to split the archive by shape into three local stores, an embedding index for prose, DuckDB for spreadsheets and a graph database for relationships, with a rule deciding where each question goes before anything is searched. Eight lookups that would have cost a million tokens of reading now cost sixteen thousand. This is what each store is good at, what picking the wrong one costs, and the measured results.

    August 25, 2026 29 min read

  • Event-time partitioning in Iceberg, and the retrieval filter that refused its own answer

    An agent reading a governed warehouse has no equivalent of a human analyst's instinct to double-check a number before repeating it, and that gap is not hypothetical. I built the reference data-and-retrieval architecture end to end on a real public feed and partitioned the lakehouse by each record's own event time rather than by the date it arrived, which is the one choice that lets a governance layer measure how unsettled a count still is. Asked an ordinary question, the retrieval pipeline's filter step used that measurement to drop two of six documents outright and label the rest with how many times each had already changed — including one drop that had been perfectly safe to cite forty-eight hours earlier.

    August 16, 2026 13 min read

  • pgvector and Neo4j sitting on top of Postgres, and the multi-hop question neither alone can answer

    A team building past prototype asked how to stand up a proprietary small language model, and the advice I gave back — try vector and graph retrieval instead, cheaper to build — undersold what was actually being decided. Each layer buys immunity from a different failure, not a discount on the same one. I grounded the claim in three systems I've actually run: a dating app's relational core in production, a graph demo that models that same app's match-and-handoff chain in Neo4j, and a governed lakehouse fusing pgvector with a lineage graph. Stacked in the right order, the three answer a question — why did a specific connection go quiet — that no single layer can answer alone, and the order you add them in is not interchangeable.

    August 16, 2026 12 min read

© 2026 Sree Dayanidhi
LinkedIn GitHub Subscribe by RSS