Two years of AI tools have converged on the same screen: one input box, one scrolling transcript, one generated document at the end. It works for quick questions. It works poorly for research, where a task lives for weeks, touches dozens of files, and deserves to be seen while it happens.
In The Premise I argued that the AI coding IDE (file tree on the left, editor in the middle, chat on the right) was accidentally the best research interface we ever had, and that we traded it away just as AI arrived. This post is the engineering follow-up: how Beeblio actually builds that workspace. It's also the first post on this blog, so it doubles as a map: the pieces that deserve their own posts (literature manager, reference manager, agent skills, cloud compute) will get them later.
One screen, three surfaces
The workspace is a full-screen application shell with three surfaces around a shared foundation:
The Beeblio workspace window: research artifact browsers on the left, the authoring system at the center with a live document, citation chips, and a results table, and the research assistant on the right proposing a cited section edit, with one cloud workspace underneath.
An activity rail on the left switches the side panel between lenses on the project: the file explorer, the literature panel, a knowledge index, skills. The center is a tabbed editor area: one tab per open file, any format. The right panel is the agent. On a phone the same three surfaces collapse into an editor with a slide-out panel and a draggable bottom sheet for the agent.
None of this layout is novel. That's the point. It's the arrangement software developers refined over decades, which researchers kept approximating with a folder window, a Word doc, and a chat tab. The work was making each surface research-native, which is the rest of this post.
Files first: conventions, not a database
The foundational decision is that a project is a file tree, stored in cloud storage, not rows in an app database:
The /workspace file tree with four protected roots: 1-References holding references.bib and papers, 2-Data with a read-only raw dataset and a derived folder, 3-Analysis with scripts and run artifacts, 4-Reports with the open manuscript.
The four root folders are numbered so sort order equals the research lifecycle, and they're protected: the app and the agent can populate them endlessly but never rename, move, or delete them. Everything inside is yours to organize however you like.
Why plain files instead of a database? Three reasons. Portability: you can download the whole project as an archive and every artifact opens with ordinary tools. Honesty: the agent's tools operate on paths, so "I cleaned the data and saved it here" is a verifiable statement about a file, not a vibe. And provenance, enforced structurally: raw datasets under 2-Data are read-only to compute jobs; cleaning, recoding, or merging must produce a new file under derived/. The original measurement is never mutated by a script that was half-right.
Markdown is the native document format: every report, memo, and draft defaults to a .md file. Office formats are export targets, not the storage layer. references.bib is the one non-Markdown exception, a real BibTeX database that the whole system treats as an API (more on that below).
The shell: optimistic state and an event bus
Most "AI workspaces" feel like web pages. This one needs to feel like an IDE, and that comes from a hundred small decisions rather than one framework choice. A few that matter most:
Tabs with real semantics. Single-clicking a file opens it as a preview, an italic tab that gets replaced by the next preview, like VS Code's single-click behavior. Double-click pins it. Tabs drag to reorder, closing a dirty tab interrupts you with save/discard, and an untitled draft routes through a Save As dialog into the file tree. Renaming or moving a file re-paths its open tab without remounting the editor, so your undo history and cursor survive. Deep links (?file=…) reopen a specific file, and ⌘P pops a Quick Open over files, library, and literature.
An event bus with optimistic updates. The panels don't share a state manager; they share a window-level event bus. A rename, move, upload, or delete broadcasts a small mutation event, and every mounted listing applies it locally and instantly, then reconciles against the server in the background. The same bus carries richer intents: an editor can hand a literature query to the Literature panel, a chat mention of a folder can ask the File Explorer to reveal it, and the agent's tool calls dispatch beeblio:agent-file-edit and beeblio:reload-workspace-file events so mounted editors update while the agent works (more on that later).
Seeded caches and prefetch. The server component embeds the project's root listing (and the text of the default file) into the page, so the first paint is synchronous and a revalidate happens in the background. Hovering a row in the explorer warms that file's content into the cache, so the eventual click opens instantly. These are small lies of preemption, and they're what make a cloud file tree feel local.
The editor: a word processor that stores Markdown
The center surface is where the classic IDE fell short: a code editor is not a word processor. Beeblio's is built on TipTap (a ProseMirror framework), with one governing rule: Markdown in, Markdown out. The visual editor and a raw source view are two windows onto the same bytes; you can flip between them mid-sentence. GFM tables, task lists, frontmatter, KaTeX math, and Mermaid diagrams all round-trip through the file on disk, which is the same file the agent reads and writes.
The research-native part is citations. A citation in the document model is a node holding a key into references.bib, persisted in Markdown as a [@key] token:
…transformer-based approaches now dominate recognition tasks
[@smith-2020-deep-learning-for-named-entity-recognition-4f2a91c7].
Tokens render as styled in-text citations, the References section is generated at the end of the document from the order citations appear, and switching citation styles re-renders every token and the bibliography in one pass. A token whose entry is missing renders as a visible "Missing reference"; the document never silently invents a source. Even the clipboard is citation-aware: copying a passage resolves its tokens into formatted citations, so pasting into an email or a slide still reads correctly.
The @ key is the editor's context menu. Typing it opens one fused search across three tiers: entries already in your bibliography (insert a citation), files in the workspace (images embed with document-relative paths; Mermaid diagrams embed as live code blocks), and (when the query is long enough) a live literature search across providers. Picking a web result saves the BibTeX entry into references.bib first, reloads the bibliography so the citation resolves, and only then replaces your @query text with the citation node. Cite-while-you-write, with the database always the source of truth.
And because it's the same cloud file tree, figures drag straight from the File Explorer into the document. No upload dialog, no asset library.
One tree, many lenses
The File Explorer is a real explorer: breadcrumbs, drag-and-drop uploads and moves, multi-select, workspace-wide search, duplicate, and zip-download. But researchers rarely think in folders; they think in artifact types. So the same file tree projects into typed views through a Research Artifact Browser:
- My Library: the bibliography database first, then literature matrices, then PDFs and bibliography files from
1-References. - Data: datasets, plus survey forms paired with their collected-responses CSV.
- Analysis: scripts, notebooks, and results from
3-Analysis. - Reports: deliverables from
4-Reports. - Figures, a grid of every chart and diagram across Analysis and Reports, with live thumbnails: images inline, Mermaid sources rendered to SVG on the spot, and HTML charts previewed in a sandboxed frame.
The important property is that these views derive from the one listing. There is no second source of truth to drift. A figure the agent drops into 3-Analysis/run-7/charts/ appears in Figures and in the explorer simultaneously, because they're the same files.
The agent: a resident with hands, not a chat widget
The right panel looks like a chat, but the architecture underneath is an agent runtime (we call it eve) speaking a typed event stream: message deltas, tool calls and results, reasoning, compaction notices. Sessions are durable: every turn is persisted with a cursor and an event snapshot, so reopening a conversation renders instantly from the snapshot and replays anything missing from the durable stream. Long sessions compact themselves mid-turn; the UI shows it as a quiet "summarizing" beat instead of a stall.
Context is structured, not pasted. Each turn begins with a workspace context block the UI assembles:
{
"files": [{ "path": "4-Reports/draft.md", "kind": "active" }],
"selections": [{ "text": "…the passage you selected…", "source": "document" }],
"interaction": { "kind": "contextual-selection", "targetFilePath": "4-Reports/draft.md" }
}
The active file is the tab you're editing (with its unsaved editor snapshot when present), so the agent never edits a stale version of what you're looking at. Selections are captured from the editor's document model, not the rendered DOM, so a selected passage contains the real [@key] tokens and can be used verbatim as an edit anchor. The interaction mode records how you invoked the agent: "insert at my caret" versus "answer about this selection" versus "transform this passage". The same sentence produces an edit or an answer depending on where you typed it.
Edits go through typed tools, not code blobs. edit_document replaces an exact, unique passage or inserts between anchors, and fails rather than guesses when the anchor is ambiguous. write_file requires the file's current generation token, an optimistic-concurrency check so a stale overwrite can't clobber a concurrent edit. references.bib rejects generic file tools entirely; citations move only through search_bibliography and update_bibliography, which is what keeps every [@key] in every document resolvable. Around these sits a full research toolbelt: literature search and paper-details lookup, PDF and Office reading, audio transcription, image analysis, web search, a vector knowledge index over project files.
Compute is disposable. Heavy work runs through one tool: run_analysis executes a command in an isolated cloud job with a pinned scientific Python stack (NumPy, pandas, Polars, SciPy, Matplotlib, Seaborn, scikit-learn, statsmodels, DuckDB, PyArrow). No network, no inherited secrets, raw data mounted read-only. The intended rhythm is: the agent writes a real script into 3-Analysis with write_file, then executes that file by path, so every result is reproducible from an artifact you can read, not a command you can't.
And you watch it happen. This is the loop that makes the three-pane layout worth building. When the agent starts a tool call that will touch 4-Reports/draft.md, the editor shows it's being edited; when the write completes, a per-file reload event updates the mounted editor in place, often before the turn finishes talking. At turn end, a workspace-changed event re-lists every panel, catching anything done through a shell. The chat doesn't hand you a document at the end. The workspace changes in front of you.
Models are routed through a hosted model gateway, and you can bring your own key and model; a BYOK turn fails closed rather than silently falling back to our system model.
What's next
This post covered the chassis. The parts I deliberately skimmed each get their own write-up soon:
- the literature manager: providers, dedup, open-access awareness, the reference sheet;
- the literature map: synthesis views over a corpus, starting with the comparison matrix;
- agent skills: how users teach the agent reusable, versioned procedures;
- the reference manager:
references.bibas an API, citation styles, export pipelines; - cloud analytical packages: the disposable compute story, and what runs where.
If the argument resonates, read The Premise for the why. This is the how. The tool is built in the open, feedback is welcome, and every milestone lands here.