Straight answers about what opti-pipe actually does, what data it touches, and what it costs right now β grouped so you can jump to the part you care about.
Product
Two things, mainly. First, the core recommendation engine is a deterministic rule engine, not a model guessing at your log β it reads structured metrics (timing, row counts, memory usage) and applies tested, unit-tested heuristics (like the target size for a shuffle partition, or what actually distinguishes a real OOM risk from a config that's just conservative). The same input always produces the same recommendation, with the same reasoning, every time. A general AI agent handed a raw Spark event log or run_results.json first has to correctly parse a genuinely complicated format, then reason about it without hallucinating a plausible-sounding but wrong threshold β a well-known failure mode when you ask a model to "just look at this and tell me what's wrong."
Second, privacy: opti-pipe only ever reads numeric metrics out of your file β never your SQL, model names, DataFrame code, or file paths (see the next question). Pasting a raw log into a general chatbot exposes all of that to the model.
There's also an optional "Ask AI" second opinion layered on top of the rule engine, for when you want a sanity check on a specific recommendation β but it's a second opinion, not the foundation, and the deterministic engine works the same with or without it.
It's an optional button next to a recommendation that asks "does the deterministic engine's recommendation actually look right given the fuller context of my file?" β a second opinion from a model, on demand, not a requirement to use opti-pipe at all. The rule engine above it stays free, deterministic, and fully unit-tested either way.
Those solve a broader problem (continuous monitoring, alerting, dashboards across your whole stack) and cost accordingly. opti-pipe is narrower on purpose: upload one real artifact your pipeline already produced, get concrete, approve-before-apply configuration recommendations back. No agent to install, no platform to run, nothing to keep paying for just to watch a graph.
No. Every recommendation is approve-before-apply β opti-pipe tells you what it found and what it would change, with the reasoning behind it, and you decide. Nothing touches your actual Spark, dbt, or Flink configuration on its own.
Privacy & data
Your uploaded pipeline data is anonymized by design β only numeric metrics (timing, row counts, memory figures, and similar) are ever read from your file. Your SQL, model names, DataFrame code, and file paths are never read, stored, or sent anywhere, including to the optional "Ask AI" feature.
No. Upload your own file (or explore the live sample data) and you get one real, full recommendation per pipeline for free, no signup required. An account unlocks every recommendation for that pipeline, not just the first.
Pricing & access
Yes, right now β opti-pipe is in early access and free to use, no card required, nothing charged. Paid plans are coming once billing is fully turned on; nothing about your account changes automatically when that happens, and existing early-access users will hear about it directly before anything is billed.
No. Signing up today doesn't collect a card and doesn't start a billing relationship of any kind β there's simply nothing to be charged on. If and when paid plans launch, that'll be a clear, separate opt-in, not something that happens silently to an existing account.
Technical
A Spark event log, dbt's target/run_results.json, or a Flink metrics export from its REST API β one genuine artifact your own run already produced, not something you fill out by hand. See the upload guide for exactly where to find each one.
Yes β the dashboard has live sample pipelines you can explore with zero setup, so you can see what a real recommendation looks like before ever touching your own files.
The fastest way to get a real answer is to just try it β upload a file or explore the sample data, free, no account needed for the first recommendation.