Files
dbt_cloud/README.md
T
2023-04-09 15:13:42 +08:00

184 lines
5.3 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# DBT governance
Thid markdown document will specify the rules that need to follow by every developer.
<br>
Short naming converntion lists that need to follow
| model_type Shortcut | Full name |
| ------------------- | ------ |
| seed | Seed |
| src | Source |
| snap | Snapshot |
| stg | Stage |
| int | Intermediate |
| fct | Fact |
| dim | Dimension |
| rep | Report |
| sem | Semantic |
| model_name Shortcut | Full name | Use case |
| ------------------- | ------ | ------ |
| brg | Brigde | relationship table |
| log | Log | log data |
## CIEF General Rule
- Snowflake SQL use to write dbt SQL code
- Star schema is use
- First week of the day is start at Monday, and last week of the day will be Sunday
- Fiscal 1st quarter is start from Febuary
- Timezone: Kuala Lumpur / Malaysia (GMT +8)
- Some status/type is not input in raw database (You should found in seed)
<br/>
## DBT Rule
### Structure of the dbt model layer
- Analyses
1. Analyses layer will not create in Snowflake
2. You can write some adhoc queries
3. Test some SQL code before write in Models
<br/>
- Macros
1. Jinja function that reuse in other models
<br/>
- Models
- Staging
- System name
1. Materialise: `View`
2. 1-to-1 relationship (or mapping) to source tables.
3. Column renaming
4. Column remove
5. Column data renaming, such as (status integer to string)
6. Data input error cleansing
7. Check for raw data freshness
8. Data type transformation
9. Data split and merge
- Source
1. Source will write in YML
2. Check Freshness
3. Check duplicate id
4. Check null id
- Intermediate
1. Materialise: `Ephemerally`
2. Stacking layers of logic with clear and specific purposes to prepa our staging models to join into the entities we want
3. Be referenced repeatedly in more than one model
4. Isolating complex operations
- Marts/warehouse
1. Materialise: `Table`
2. Store fact and dimension models
- Marts/reporting
1. Materialise: `Table`
2. Store custom reports
<br/>
- Seeds
1. Store custom dataset in csv format
2. The dataset must not change frequently
<br/>
- Snapshots
1. DBT build-in SCD type 2 function
2. Prefer to use after the source table and before the staging layer
<br/>
- Tests
1. Can write some yml test case
<br/>
### File Naming Rules
- File names must be unique
- Each sql/python file must start with [`model_type`]_[`system_name`]__[`model_name`]s.sql/py
- The model should be name as plural. (eg: stg_exchange__orders.sql)
<br/>
### Column Naming Rules
- If array datatype, the column naming must be plural (eg: ids, messages)
- Schema, table and column names should be in `snake_case`.
- Limit use of abbreviations that are related to domain knowledge. An onboarding
employee will understand `current_order_status` better than `current_os`.
- Use names based on the _business_ terminology, rather than the source terminology.
- Each model should have a primary key that can identify the unique row, and should be named `<object>_id`, e.g. `account_id` this makes it easier to know what `id` is being referenced in downstream joined models.
- If a surrogate key is created, it should be named `<object>_sk`.
- For `base` or `staging` models, columns should be ordered in categories, where identifiers are first and date/time fields are at the end.
Example:
```sql
transformed as (
select
-- ids
order_id,
customer_id,
-- dimensions
order_status,
is_shipped,
-- measures
order_total,
-- date/times
created_at,
updated_at,
-- metadata
_sdc_batched_at
from source
)
```
- Date/time columns should be named according to these conventions:
- Timestamps: `<event>_datetime`
Example: `created_datetime`
- Dates: `<event>_date`
Example: `created_date`
- Booleans should be prefixed with `is_` or `has_`.
Example: `is_active_customer` and `has_admin_access`
- Price/revenue fields should be in decimal currency (e.g. `19.99` for $19.99; many app databases store prices as integers in cents). If non-decimal currency is used, indicate this with suffix, e.g. `price_in_cents`.
<br/>
### SQL Coding Rule
- The SQL clause **must** be UPPERCASE (eg. `SELECT`, `FROM`, `WHERE`, `GROUP BY`, `LIMIT`, `WITH`, `AS`, `SUM`, `PARTITION OVER`, `DIV0`, `LEFT JOIN`, `ON` ....)
- Try to avoid using `EXCLUDE` in the SQL
- Avoid using multi layer subquery, and try to use CTE, subquery must not more than 1 layer
- Each model structure must contain with
```sql
--IMPORT
WITH [table_names] AS (
SELECT * FROM {{ ref('file_names')}}
),
--LOGIC
[logic_names] AS (
.....
),
--FINAL
final__[table_names] AS (
)
SELECT * FROM final__[final_table_names]
```
### Testing
- At a minimum, `unique` and `not_null` tests should be applied to the expected primary key of each model.