Pipeline run
73c8738c-c4c0-45e0-8da9-b126e86c3f61
Client output enrichment
v2 Skill cluster · Nature of work · AI index · Tech stack maturity · Evidence · KRA descriptionvocab breakdown (legacy)
1 LLM mention extraction + deterministic catalog sweep
2 Six-layer resolution cascade (exact → alias → fuzzy → embedding → judge → AI-prediction lane)
3 Table-lookup layering: dim_members → role_dim_map tiers, served verbatim
Data Engineer
in_db name family · data-engineerslug: data-engineer · role_id: 1839a54d-e909-543e-b36e-745551d4e000 · catalog: skill_library_v4 (tiers served verbatim from role_dim_map)
The company is seeking a Data Engineer with over 5 years of experience to join their analytics platform team. The role requires proficiency in building data pipelines using Python and SQL, along with hands-on experience with Apache Spark and Apache Airflow for orchestration. Distinctive requirements include strong Unix shell scripting skills for automation and experience with Kafka for streaming ingestion. Familiarity with CP4D for governed data workloads and a cloud data warehouse like Snowflake is also necessary.
Job description
Job Title: Data Engineer We are hiring a Data Engineer for our analytics platform team. Requirements: - 5+ years building data pipelines with Python and SQL - Hands-on experience with Apache Spark and Apache Airflow for orchestration - Experience with CP4D for governed data workloads - Strong Unix shell scripting for automation - Experience with Kafka for streaming ingestion - Snowflake or another cloud data warehouse - Git-based workflows and CI discipline
Skills from this JD
L0–L3 tiers come verbatim from the role's role_dim_map; violet AI GENERATED chips are AI-prediction-lane skills pending review in the queue.
This skill is in Secondary as the JD says "Experience with CP4D for governed data workloads" - it isn't in this role's skill catalog yet, so it's tracked for review.
This skill is in Secondary as the JD says "Strong Unix shell scripting for automation" - it isn't in this role's skill catalog yet, so it's tracked for review.
Apache Beam — Data Processing & Pipeline Frameworks (L0)
Scala — Programming Languages (L0)
Prefect — Workflow Orchestration (L0)
Google Cloud Storage — Cloud Platforms (L1)
Airbyte — Data Ingestion & Integration (L1)
Apache Parquet — Data Lake & Storage Formats (L1)
Library artifacts (this run)
nano JD Parser — gpt-4.1-nano click to toggle
Show raw JSON
{
"JD_type": "pass",
"about_company": null,
"ai_kras": [],
"certifications": [],
"client_details": null,
"company_name": null,
"company_size": null,
"ctc": null,
"domain": {
"primary": {
"aliases": [],
"domain": "Other"
},
"secondary": null
},
"education": [],
"experience": {
"max": null,
"min": 5,
"raw": "5+ years building data pipelines with Python and SQL"
},
"job_locations": [],
"notice_period": {
"days": null,
"raw": null
},
"open_to_relocate": false,
"role": "Data Engineer",
"role_aliases": [
{
"name": "Data Engineer",
"reasoning": "generalized form of the picked role",
"relation": "synonym"
}
],
"role_archetype": "Data",
"roles_and_responsibilities": [
{
"bullet_count": 7,
"heading": "Requirements",
"heading_was_present": true,
"source_marker": {
"first_5_words": "Requirements: - 5+ years building",
"last_5_words": "workflows and CI discipline"
},
"text": "- 5+ years building data pipelines with Python and SQL\n- Hands-on experience with Apache Spark and Apache Airflow for orchestration\n- Experience with CP4D for governed data workloads\n- Strong Unix shell scripting for automation\n- Experience with Kafka for streaming ingestion\n- Snowflake or another cloud data warehouse\n- Git-based workflows and CI discipline",
"word_count": 43
}
],
"urls": []
}
API 1 — extract-from-jd click to toggle
{
"catalog": "v4",
"consider_adding": [
{
"dimension": "Data Processing \u0026 Pipeline Frameworks",
"reason": "In the role canon\u0027s Data Processing \u0026 Pipeline Frameworks - adding it narrows the candidate pool.",
"skill": "Apache Beam",
"tier": "L0"
},
{
"dimension": "Programming Languages",
"reason": "In the role canon\u0027s Programming Languages - adding it narrows the candidate pool.",
"skill": "Scala",
"tier": "L0"
},
{
"dimension": "Workflow Orchestration",
"reason": "In the role canon\u0027s Workflow Orchestration - adding it narrows the candidate pool.",
"skill": "Prefect",
"tier": "L0"
},
{
"dimension": "Cloud Platforms",
"reason": "In the role canon\u0027s Cloud Platforms - adding it narrows the candidate pool.",
"skill": "Google Cloud Storage",
"tier": "L1"
},
{
"dimension": "Data Ingestion \u0026 Integration",
"reason": "In the role canon\u0027s Data Ingestion \u0026 Integration - adding it narrows the candidate pool.",
"skill": "Airbyte",
"tier": "L1"
},
{
"dimension": "Data Lake \u0026 Storage Formats",
"reason": "In the role canon\u0027s Data Lake \u0026 Storage Formats - adding it narrows the candidate pool.",
"skill": "Apache Parquet",
"tier": "L1"
}
],
"final_skills": [
{
"dimension": {
"display_name": "Programming Languages",
"slug": "programming-languages"
},
"is_primary": true,
"layer": "L0",
"layer_source": "catalog_v4",
"origin": "catalog",
"rationale": "the craft is exercised in Python and SQL; every owned artifact \u2014 transform, DAG, quality check \u2014 is authored in them.",
"skill_id": "646cb350-7c71-5ed2-971e-d141c566ce6e",
"skill_name": "Python"
},
{
"dimension": {
"display_name": "Programming Languages",
"slug": "programming-languages"
},
"is_primary": true,
"layer": "L0",
"layer_source": "catalog_v4",
"origin": "catalog",
"rationale": "the craft is exercised in Python and SQL; every owned artifact \u2014 transform, DAG, quality check \u2014 is authored in them.",
"skill_id": "dd6b38b8-6e50-5e89-a827-5b03688e6113",
"skill_name": "SQL"
},
{
"dimension": {
"display_name": "Data Processing \u0026 Pipeline Frameworks",
"slug": "data-processing-pipeline-frameworks"
},
"is_primary": true,
"layer": "L0",
"layer_source": "catalog_v4",
"origin": "catalog",
"rationale": "transformation engines are where pipeline work physically happens; Spark and dataframe fluency is the single strongest resume signal for the",
"skill_id": "f885334e-9062-5967-bc62-b5b5a27664cf",
"skill_name": "Apache Spark"
},
{
"dimension": {
"display_name": "Workflow Orchestration",
"slug": "workflow-orchestration"
},
"is_primary": true,
"layer": "L0",
"layer_source": "catalog_v4",
"origin": "catalog",
"rationale": "the xlsx rationale names orchestration (Airflow, Dagster, Prefect) as the family\u0027s defining tooling; DAG ownership \u2014 scheduling, retries, ba",
"skill_id": "dc7c7803-0573-5cde-b2c1-a7cacccb3b6f",
"skill_name": "Apache Airflow"
},
{
"dimension": {
"display_name": "Data Ingestion \u0026 Integration",
"slug": "data-ingestion-integration"
},
"is_primary": true,
"layer": "L1",
"layer_source": "catalog_v4",
"origin": "catalog",
"rationale": "Charter owns_4 makes ingestion an owned surface: connector EL, CDC, and API extraction are how data enters everything else this role builds;",
"skill_id": "a2275eea-20e3-59dd-b201-af76b7ae706f",
"skill_name": "Apache Kafka"
},
{
"dimension": {
"display_name": "Data Warehouses \u0026 Query Engines",
"slug": "data-warehouses-query-engines"
},
"is_primary": true,
"layer": "L1",
"layer_source": "catalog_v4",
"origin": "catalog",
"rationale": "Loading and querying warehouses is daily work \u2014 owns_1\u0027s conformed outputs land there and collaborates_3 names the partnership \u2014 but platfor",
"skill_id": "de72e911-d96b-54ca-8acf-527eb58cddac",
"skill_name": "Snowflake"
},
{
"dimension": {
"display_name": "Version Control",
"slug": "version-control"
},
"is_primary": false,
"layer": "L3",
"layer_source": "catalog_v4",
"origin": "catalog",
"rationale": "Hygiene: all pipeline and DAG code lives in Git with PR review; universal expectation, weak differentiator (same tiering rationale as both p",
"skill_id": "69a9e960-a4b1-5e7e-b55c-04980465e177",
"skill_name": "Git"
},
{
"dimension": {
"display_name": "Programming Languages",
"slug": "programming-languages"
},
"is_primary": true,
"layer": "L0",
"layer_source": "catalog_v4",
"origin": "catalog",
"rationale": "the craft is exercised in Python and SQL; every owned artifact \u2014 transform, DAG, quality check \u2014 is authored in them.",
"skill_id": "9b100b65-d396-5927-85c5-cfce46ca7c77",
"skill_name": "Shell Scripting"
}
],
"history_run_id": null,
"jd_parameters": {
"certifications": [],
"clientDetails": null,
"company": null,
"companySize": null,
"ctc": {
"currency": null,
"max": null,
"min": null,
"period": null,
"raw": null
},
"educationRequirements": [],
"experience": {
"max": null,
"min": 5,
"raw": "5+ years building data pipelines with Python and SQL"
},
"industryDomain": "Other",
"knockouts": {
"certifications": [],
"ctc": {
"currency": null,
"max": null,
"min": null
},
"educationRequirements": [],
"experience": {
"max": null,
"min": 5
},
"location": [],
"noticePeriod": null
},
"locations": [],
"noticePeriod": null,
"openToRelocate": false,
"role": "Data Engineer",
"roleSynonyms": [
"Data Engineer"
]
},
"jd_summary": "The company is seeking a Data Engineer with over 5 years of experience to join their analytics platform team. The role requires proficiency in building data pipelines using Python and SQL, along with hands-on experience with Apache Spark and Apache Airflow for orchestration. Distinctive requirements include strong Unix shell scripting skills for automation and experience with Kafka for streaming ingestion. Familiarity with CP4D for governed data workloads and a cloud data warehouse like Snowflake is also necessary.",
"layer_conflicts": [],
"nano_parsed": {
"JD_type": "pass",
"about_company": null,
"ai_kras": [],
"certifications": [],
"client_details": null,
"company_name": null,
"company_size": null,
"ctc": null,
"domain": {
"primary": {
"aliases": [],
"domain": "Other"
},
"secondary": null
},
"education": [],
"experience": {
"max": null,
"min": 5,
"raw": "5+ years building data pipelines with Python and SQL"
},
"job_locations": [],
"notice_period": {
"days": null,
"raw": null
},
"open_to_relocate": false,
"role": "Data Engineer",
"role_aliases": [
{
"name": "Data Engineer",
"reasoning": "generalized form of the picked role",
"relation": "synonym"
}
],
"role_archetype": "Data",
"roles_and_responsibilities": [
{
"bullet_count": 7,
"heading": "Requirements",
"heading_was_present": true,
"source_marker": {
"first_5_words": "Requirements: - 5+ years building",
"last_5_words": "workflows and CI discipline"
},
"text": "- 5+ years building data pipelines with Python and SQL\n- Hands-on experience with Apache Spark and Apache Airflow for orchestration\n- Experience with CP4D for governed data workloads\n- Strong Unix shell scripting for automation\n- Experience with Kafka for streaming ingestion\n- Snowflake or another cloud data warehouse\n- Git-based workflows and CI discipline",
"word_count": 43
}
],
"urls": []
},
"pipeline": "v4",
"rejected": false,
"rejection_code": null,
"rejection_reason": null,
"role": {
"canonical_name": "Data Engineer",
"family": "data-engineer",
"match_method": "name",
"resolution": "in_db",
"role_id": "1839a54d-e909-543e-b36e-745551d4e000",
"similarity": null,
"slug": "data-engineer"
},
"run_id": "jdv4-0d3b5c6c78e4",
"secondary_meta": [
{
"audit_reasoning": "This skill is in Secondary as the JD says \"Experience with CP4D for governed data workloads\" - it isn\u0027t in this role\u0027s skill catalog yet, so it\u0027s tracked for review.",
"jd_quote": "Experience with CP4D for governed data workloads",
"origin": "ai_predicted",
"provenance": "verb-backed",
"reason_code": "not_in_catalog",
"skill": "IBM Cloud Pak for Data"
},
{
"audit_reasoning": "This skill is in Secondary as the JD says \"Strong Unix shell scripting for automation\" - it isn\u0027t in this role\u0027s skill catalog yet, so it\u0027s tracked for review.",
"jd_quote": "Strong Unix shell scripting for automation",
"origin": "ai_predicted",
"provenance": "verb-backed",
"reason_code": "not_in_catalog",
"skill": "Unix Shell Scripting"
}
],
"secondary_skills": [
"IBM Cloud Pak for Data",
"Unix Shell Scripting"
],
"skill_layers": [
{
"label": "Anchor",
"layer": "L0",
"skills": [
{
"dimension": {
"display_name": "Programming Languages",
"slug": "programming-languages"
},
"name": "Python",
"origin": "catalog",
"rationale": "the craft is exercised in Python and SQL; every owned artifact \u2014 transform, DAG, quality check \u2014 is authored in them."
},
{
"dimension": {
"display_name": "Programming Languages",
"slug": "programming-languages"
},
"name": "SQL",
"origin": "catalog",
"rationale": "the craft is exercised in Python and SQL; every owned artifact \u2014 transform, DAG, quality check \u2014 is authored in them."
},
{
"dimension": {
"display_name": "Data Processing \u0026 Pipeline Frameworks",
"slug": "data-processing-pipeline-frameworks"
},
"name": "Apache Spark",
"origin": "catalog",
"rationale": "transformation engines are where pipeline work physically happens; Spark and dataframe fluency is the single strongest resume signal for the"
},
{
"dimension": {
"display_name": "Workflow Orchestration",
"slug": "workflow-orchestration"
},
"name": "Apache Airflow",
"origin": "catalog",
"rationale": "the xlsx rationale names orchestration (Airflow, Dagster, Prefect) as the family\u0027s defining tooling; DAG ownership \u2014 scheduling, retries, ba"
},
{
"dimension": {
"display_name": "Programming Languages",
"slug": "programming-languages"
},
"name": "Shell Scripting",
"origin": "catalog",
"rationale": "the craft is exercised in Python and SQL; every owned artifact \u2014 transform, DAG, quality check \u2014 is authored in them."
}
]
},
{
"label": "Primary",
"layer": "L1",
"skills": [
{
"dimension": {
"display_name": "Data Ingestion \u0026 Integration",
"slug": "data-ingestion-integration"
},
"name": "Apache Kafka",
"origin": "catalog",
"rationale": "Charter owns_4 makes ingestion an owned surface: connector EL, CDC, and API extraction are how data enters everything else this role builds;"
},
{
"dimension": {
"display_name": "Data Warehouses \u0026 Query Engines",
"slug": "data-warehouses-query-engines"
},
"name": "Snowflake",
"origin": "catalog",
"rationale": "Loading and querying warehouses is daily work \u2014 owns_1\u0027s conformed outputs land there and collaborates_3 names the partnership \u2014 but platfor"
}
]
},
{
"label": "Hygiene",
"layer": "L3",
"skills": [
{
"dimension": {
"display_name": "Version Control",
"slug": "version-control"
},
"name": "Git",
"origin": "catalog",
"rationale": "Hygiene: all pipeline and DAG code lives in Git with PR review; universal expectation, weak differentiator (same tiering rationale as both p"
}
]
}
],
"unmapped_skills": []
}
API 2 — extract-details
{}
API 3 — final-role-output
{}
LLM Calls
Every model call made for this run, in pipeline order. Click a card to see the model's response.