Semantic layer: pre-enrich or blank slate
When you connect a database, edisyl does two very different things to it. The first is discovery: the connector lists the databases, schemas, tables, and columns the connecting user can see and imports them as models. This is deterministic, fast, and spends nothing on an LLM. The result is a catalog of names.
The second is enrichment, and it’s the step that makes a source answerable. A study runs over the imported tables and learns what the columns mean, how the tables join, which measures and dimensions they support, and which tables a given question is actually about. That learned understanding is the semantic layer. Until it exists for a table, the pack knows the table’s name and nothing else. A question that depends on it is declined as though the data were never connected.
The last step of the Connect a source wizard asks how you want that second step to happen. The two options differ only in when and how much enrichment runs, not in what gets connected.
Pre-enrich models
Section titled “Pre-enrich models”Every table you selected is studied up front, in one pass, as soon as the source connects. The study proposes joins between tables, measures and dimensions on them, and example questions the source can answer. Those proposals land for you to review.
What you get: a source that can answer questions about any of its selected tables from the first ask, and a set of proposals that shows you what the pack thinks the data means before anyone relies on it.
What it costs: one LLM-driven study per table, paid immediately, including for tables nobody ends up asking about.
Start with a blank slate
Section titled “Start with a blank slate”Only the schema is indexed. No table is studied until there’s a reason to. The semantic layer then grows in response to use: when someone asks a question the pack can’t ground in modeled data, that’s a data-miss, and it surfaces as a learning candidate, a concrete piece of modeling the pack has worked out it needs. You close it by tasking the pack to curate that part of the layer (see Closing gaps: teach vs. task), and the next ask lands.
What you get: enrichment spent only on tables that real questions touch, in the order people actually need them, and a layer whose shape reflects how your team uses the data rather than everything the schema happens to contain.
What it costs: the first ask against any unstudied table is declined, and someone has to route the gap before it can be answered.
Choosing between them
Section titled “Choosing between them”| Pre-enrich models | Start with a blank slate | |
|---|---|---|
| First ask on a table | Answers | Declined until curated |
| Enrichment spend | All selected tables, up front | Only tables that asks reach, over time |
| Review work | Proposals to review at connect time | Learning candidates to act on as they surface |
| Layer reflects | The schema you selected | The questions your team asks |
Pick pre-enrich when
- You already narrowed the source to the tables that matter, on the Database and Tables steps. A tight selection is cheap to study in full and there’s little waste.
- You know the questions coming and want them answerable on day one, for a demo, an onboarding, or a team that won’t tolerate “not yet” on its first try.
- You want to see the pack’s proposed joins and measures before the data is used, so you can correct them early.
Pick a blank slate when
- The schema is wide and most of it is noise. A warehouse with hundreds of tables where a dozen carry the questions is the classic case: pre-enriching all of it buys understanding nobody will use.
- You don’t yet know what people will ask. Letting the layer grow from real questions tells you which tables matter and avoids modeling the rest.
- You want to control spend and grow it deliberately, curating one gap at a time.
- The schema is still changing. Enriching tables that will be renamed or dropped next week is wasted work.
The choice isn’t final
Section titled “The choice isn’t final”A blank slate source can be enriched in full at any time, and a pre-enriched source needs re-enrichment when its schema changes. Both routes end up in the same place, a semantic layer over the tables you care about.
- Blank slate, later enriched in bulk. Run
edisyl data-sources enrich <id>on the source, or task the pack with a broader curation objective. This studies every imported table, exactly as pre-enrich would have. - New tables added to the warehouse. They aren’t catalogued automatically, and
enrichwould re-study every existing model too. Re-catalog withimport --no-study, thenstudyonly the new models. Adding tables to a connected source walks it through, including the wait between a study finishing and its results being applied. - Either route, one gap at a time.
edisyl task <pack> "..."with a specific objective curates the part of the layer that objective needs, whichever starting point the source began with.
How this maps to the CLI
Section titled “How this maps to the CLI”The CLI doesn’t offer the choice. data-sources create runs discovery and an import that studies every discovered model by default, which is the pre-enrich path. There’s no flag on create to skip the study, so a blank slate start is only available from the web app today.
# Discovery, import, and a study of every table, in one callDS=$(edisyl data-sources create --name "Warehouse Production" --type snowflake \ --config-file sf.json --database PROD -j | jq -r '.id')
# Later, after a schema change: re-study every modeledisyl data-sources enrich "$DS"Studies land as learning candidates that a separate review step applies, so allow some minutes before the first ask. See Connect a data source for the full flow, Adding tables to a connected source for what happens after, and Data sources for the command reference.
Was this page helpful?
Thanks for the feedback.
