Pipeline run
b7328f64-2252-4855-8fad-1d5a4c31ec1a
Client output enrichment
v2 Skill cluster · Nature of work · AI index · Tech stack maturity · Evidence · KRA descriptionvocab breakdown (legacy)
1 LLM mention extraction + deterministic catalog sweep
2 Six-layer resolution cascade (exact → alias → fuzzy → embedding → judge → AI-prediction lane)
3 Table-lookup layering: dim_members → role_dim_map tiers, served verbatim
Data Engineer
in_db name family · data-engineerslug: data-engineer · role_id: 1839a54d-e909-543e-b36e-745551d4e000 · catalog: skill_library_v4 (tiers served verbatim from role_dim_map)
We are seeking a Data Engineer with over 5 years of experience for our analytics platform team. The role requires proficiency in building data pipelines using Python and SQL, along with hands-on experience in Apache Spark and Apache Airflow for orchestration. Candidates should also have experience with Kafka for streaming ingestion and familiarity with cloud data warehouses like Snowflake. Strong Unix shell scripting skills for automation and a background in Git-based workflows are essential.
Job description
Job Title: Data Engineer We are hiring a Data Engineer for our analytics platform team. Requirements: - 5+ years building data pipelines with Python and SQL - Hands-on experience with Apache Spark and Apache Airflow for orchestration - Experience with CP4D for governed data workloads - Strong Unix shell scripting for automation - Experience with Kafka for streaming ingestion - Snowflake or another cloud data warehouse - Git-based workflows and CI discipline
Skills from this JD
L0–L3 tiers come verbatim from the role's role_dim_map; violet AI GENERATED chips are AI-prediction-lane skills pending review in the queue.
Apache Beam — Data Processing & Pipeline Frameworks (L0)
Scala — Programming Languages (L0)
Prefect — Workflow Orchestration (L0)
Google Cloud Storage — Cloud Platforms (L1)
Airbyte — Data Ingestion & Integration (L1)
Apache Parquet — Data Lake & Storage Formats (L1)
Library artifacts (this run)
nano JD Parser — gpt-4.1-nano click to toggle
Show raw JSON
{
"JD_type": "pass",
"about_company": null,
"ai_kras": [],
"certifications": [],
"client_details": null,
"company_name": null,
"company_size": null,
"ctc": null,
"domain": {
"primary": {
"aliases": [],
"domain": "Other"
},
"secondary": null
},
"education": [],
"experience": {
"max": null,
"min": 5,
"raw": "5+ years building data pipelines with Python and SQL"
},
"job_locations": [],
"notice_period": {
"days": null,
"raw": null
},
"open_to_relocate": false,
"role": "Data Engineer",
"role_aliases": [
{
"name": "Data Engineer",
"reasoning": "generalized form of the picked role",
"relation": "synonym"
}
],
"role_archetype": "Data",
"roles_and_responsibilities": [
{
"bullet_count": 7,
"heading": "Requirements",
"heading_was_present": true,
"source_marker": {
"first_5_words": "Requirements: - 5+ years building",
"last_5_words": "workflows and CI discipline"
},
"text": "- 5+ years building data pipelines with Python and SQL\n- Hands-on experience with Apache Spark and Apache Airflow for orchestration\n- Experience with CP4D for governed data workloads\n- Strong Unix shell scripting for automation\n- Experience with Kafka for streaming ingestion\n- Snowflake or another cloud data warehouse\n- Git-based workflows and CI discipline",
"word_count": 45
}
],
"urls": []
}
API 1 — extract-from-jd click to toggle
{
"catalog": "v4",
"consider_adding": [
{
"dimension": "Data Processing \u0026 Pipeline Frameworks",
"reason": "In the role canon\u0027s Data Processing \u0026 Pipeline Frameworks - adding it narrows the candidate pool.",
"skill": "Apache Beam",
"tier": "L0"
},
{
"dimension": "Programming Languages",
"reason": "In the role canon\u0027s Programming Languages - adding it narrows the candidate pool.",
"skill": "Scala",
"tier": "L0"
},
{
"dimension": "Workflow Orchestration",
"reason": "In the role canon\u0027s Workflow Orchestration - adding it narrows the candidate pool.",
"skill": "Prefect",
"tier": "L0"
},
{
"dimension": "Cloud Platforms",
"reason": "In the role canon\u0027s Cloud Platforms - adding it narrows the candidate pool.",
"skill": "Google Cloud Storage",
"tier": "L1"
},
{
"dimension": "Data Ingestion \u0026 Integration",
"reason": "In the role canon\u0027s Data Ingestion \u0026 Integration - adding it narrows the candidate pool.",
"skill": "Airbyte",
"tier": "L1"
},
{
"dimension": "Data Lake \u0026 Storage Formats",
"reason": "In the role canon\u0027s Data Lake \u0026 Storage Formats - adding it narrows the candidate pool.",
"skill": "Apache Parquet",
"tier": "L1"
}
],
"final_skills": [
{
"dimension": {
"display_name": "Programming Languages",
"slug": "programming-languages"
},
"is_primary": true,
"layer": "L0",
"layer_source": "catalog_v4",
"origin": "catalog",
"rationale": "the craft is exercised in Python and SQL; every owned artifact \u2014 transform, DAG, quality check \u2014 is authored in them.",
"skill_id": "646cb350-7c71-5ed2-971e-d141c566ce6e",
"skill_name": "Python"
},
{
"dimension": {
"display_name": "Programming Languages",
"slug": "programming-languages"
},
"is_primary": true,
"layer": "L0",
"layer_source": "catalog_v4",
"origin": "catalog",
"rationale": "the craft is exercised in Python and SQL; every owned artifact \u2014 transform, DAG, quality check \u2014 is authored in them.",
"skill_id": "dd6b38b8-6e50-5e89-a827-5b03688e6113",
"skill_name": "SQL"
},
{
"dimension": {
"display_name": "Data Processing \u0026 Pipeline Frameworks",
"slug": "data-processing-pipeline-frameworks"
},
"is_primary": true,
"layer": "L0",
"layer_source": "catalog_v4",
"origin": "catalog",
"rationale": "transformation engines are where pipeline work physically happens; Spark and dataframe fluency is the single strongest resume signal for the",
"skill_id": "f885334e-9062-5967-bc62-b5b5a27664cf",
"skill_name": "Apache Spark"
},
{
"dimension": {
"display_name": "Workflow Orchestration",
"slug": "workflow-orchestration"
},
"is_primary": true,
"layer": "L0",
"layer_source": "catalog_v4",
"origin": "catalog",
"rationale": "the xlsx rationale names orchestration (Airflow, Dagster, Prefect) as the family\u0027s defining tooling; DAG ownership \u2014 scheduling, retries, ba",
"skill_id": "dc7c7803-0573-5cde-b2c1-a7cacccb3b6f",
"skill_name": "Apache Airflow"
},
{
"dimension": {
"display_name": "Cloud Platforms",
"slug": "cloud-platforms"
},
"is_primary": true,
"layer": "L1",
"layer_source": "catalog_v4",
"origin": "catalog",
"rationale": "The modern data stack is cloud-native end to end: object stores, managed warehouses, serverless ETL, and IAM are the default substrate of ev",
"skill_id": "4629223b-c2e1-5282-9955-e7f2f8b5e07e",
"skill_name": "IBM Cloud Pak for Data"
},
{
"dimension": {
"display_name": "Data Ingestion \u0026 Integration",
"slug": "data-ingestion-integration"
},
"is_primary": true,
"layer": "L1",
"layer_source": "catalog_v4",
"origin": "catalog",
"rationale": "Charter owns_4 makes ingestion an owned surface: connector EL, CDC, and API extraction are how data enters everything else this role builds;",
"skill_id": "a2275eea-20e3-59dd-b201-af76b7ae706f",
"skill_name": "Apache Kafka"
},
{
"dimension": {
"display_name": "Data Warehouses \u0026 Query Engines",
"slug": "data-warehouses-query-engines"
},
"is_primary": true,
"layer": "L1",
"layer_source": "catalog_v4",
"origin": "catalog",
"rationale": "Loading and querying warehouses is daily work \u2014 owns_1\u0027s conformed outputs land there and collaborates_3 names the partnership \u2014 but platfor",
"skill_id": "de72e911-d96b-54ca-8acf-527eb58cddac",
"skill_name": "Snowflake"
},
{
"dimension": {
"display_name": "Version Control",
"slug": "version-control"
},
"is_primary": false,
"layer": "L3",
"layer_source": "catalog_v4",
"origin": "catalog",
"rationale": "Hygiene: all pipeline and DAG code lives in Git with PR review; universal expectation, weak differentiator (same tiering rationale as both p",
"skill_id": "69a9e960-a4b1-5e7e-b55c-04980465e177",
"skill_name": "Git"
},
{
"dimension": {
"display_name": "Programming Languages",
"slug": "programming-languages"
},
"is_primary": true,
"layer": "L0",
"layer_source": "catalog_v4",
"origin": "catalog",
"rationale": "the craft is exercised in Python and SQL; every owned artifact \u2014 transform, DAG, quality check \u2014 is authored in them.",
"skill_id": "9b100b65-d396-5927-85c5-cfce46ca7c77",
"skill_name": "Shell Scripting"
}
],
"history_run_id": null,
"jd_parameters": {
"certifications": [],
"clientDetails": null,
"company": null,
"companySize": null,
"ctc": {
"currency": null,
"max": null,
"min": null,
"period": null,
"raw": null
},
"educationRequirements": [],
"experience": {
"max": null,
"min": 5,
"raw": "5+ years building data pipelines with Python and SQL"
},
"industryDomain": "Other",
"knockouts": {
"certifications": [],
"ctc": {
"currency": null,
"max": null,
"min": null
},
"educationRequirements": [],
"experience": {
"max": null,
"min": 5
},
"location": [],
"noticePeriod": null
},
"locations": [],
"noticePeriod": null,
"openToRelocate": false,
"role": "Data Engineer",
"roleSynonyms": [
"Data Engineer"
]
},
"jd_summary": "We are seeking a Data Engineer with over 5 years of experience for our analytics platform team. The role requires proficiency in building data pipelines using Python and SQL, along with hands-on experience in Apache Spark and Apache Airflow for orchestration. Candidates should also have experience with Kafka for streaming ingestion and familiarity with cloud data warehouses like Snowflake. Strong Unix shell scripting skills for automation and a background in Git-based workflows are essential.",
"layer_conflicts": [],
"nano_parsed": {
"JD_type": "pass",
"about_company": null,
"ai_kras": [],
"certifications": [],
"client_details": null,
"company_name": null,
"company_size": null,
"ctc": null,
"domain": {
"primary": {
"aliases": [],
"domain": "Other"
},
"secondary": null
},
"education": [],
"experience": {
"max": null,
"min": 5,
"raw": "5+ years building data pipelines with Python and SQL"
},
"job_locations": [],
"notice_period": {
"days": null,
"raw": null
},
"open_to_relocate": false,
"role": "Data Engineer",
"role_aliases": [
{
"name": "Data Engineer",
"reasoning": "generalized form of the picked role",
"relation": "synonym"
}
],
"role_archetype": "Data",
"roles_and_responsibilities": [
{
"bullet_count": 7,
"heading": "Requirements",
"heading_was_present": true,
"source_marker": {
"first_5_words": "Requirements: - 5+ years building",
"last_5_words": "workflows and CI discipline"
},
"text": "- 5+ years building data pipelines with Python and SQL\n- Hands-on experience with Apache Spark and Apache Airflow for orchestration\n- Experience with CP4D for governed data workloads\n- Strong Unix shell scripting for automation\n- Experience with Kafka for streaming ingestion\n- Snowflake or another cloud data warehouse\n- Git-based workflows and CI discipline",
"word_count": 45
}
],
"urls": []
},
"pipeline": "v4",
"rejected": false,
"rejection_code": null,
"rejection_reason": null,
"role": {
"canonical_name": "Data Engineer",
"family": "data-engineer",
"match_method": "name",
"resolution": "in_db",
"role_id": "1839a54d-e909-543e-b36e-745551d4e000",
"similarity": null,
"slug": "data-engineer"
},
"run_id": "jdv4-dfb3f7c56ec5",
"secondary_meta": [],
"secondary_skills": [],
"skill_layers": [
{
"label": "Anchor",
"layer": "L0",
"skills": [
{
"dimension": {
"display_name": "Programming Languages",
"slug": "programming-languages"
},
"name": "Python",
"origin": "catalog",
"rationale": "the craft is exercised in Python and SQL; every owned artifact \u2014 transform, DAG, quality check \u2014 is authored in them."
},
{
"dimension": {
"display_name": "Programming Languages",
"slug": "programming-languages"
},
"name": "SQL",
"origin": "catalog",
"rationale": "the craft is exercised in Python and SQL; every owned artifact \u2014 transform, DAG, quality check \u2014 is authored in them."
},
{
"dimension": {
"display_name": "Data Processing \u0026 Pipeline Frameworks",
"slug": "data-processing-pipeline-frameworks"
},
"name": "Apache Spark",
"origin": "catalog",
"rationale": "transformation engines are where pipeline work physically happens; Spark and dataframe fluency is the single strongest resume signal for the"
},
{
"dimension": {
"display_name": "Workflow Orchestration",
"slug": "workflow-orchestration"
},
"name": "Apache Airflow",
"origin": "catalog",
"rationale": "the xlsx rationale names orchestration (Airflow, Dagster, Prefect) as the family\u0027s defining tooling; DAG ownership \u2014 scheduling, retries, ba"
},
{
"dimension": {
"display_name": "Programming Languages",
"slug": "programming-languages"
},
"name": "Shell Scripting",
"origin": "catalog",
"rationale": "the craft is exercised in Python and SQL; every owned artifact \u2014 transform, DAG, quality check \u2014 is authored in them."
}
]
},
{
"label": "Primary",
"layer": "L1",
"skills": [
{
"dimension": {
"display_name": "Cloud Platforms",
"slug": "cloud-platforms"
},
"name": "IBM Cloud Pak for Data",
"origin": "catalog",
"rationale": "The modern data stack is cloud-native end to end: object stores, managed warehouses, serverless ETL, and IAM are the default substrate of ev"
},
{
"dimension": {
"display_name": "Data Ingestion \u0026 Integration",
"slug": "data-ingestion-integration"
},
"name": "Apache Kafka",
"origin": "catalog",
"rationale": "Charter owns_4 makes ingestion an owned surface: connector EL, CDC, and API extraction are how data enters everything else this role builds;"
},
{
"dimension": {
"display_name": "Data Warehouses \u0026 Query Engines",
"slug": "data-warehouses-query-engines"
},
"name": "Snowflake",
"origin": "catalog",
"rationale": "Loading and querying warehouses is daily work \u2014 owns_1\u0027s conformed outputs land there and collaborates_3 names the partnership \u2014 but platfor"
}
]
},
{
"label": "Hygiene",
"layer": "L3",
"skills": [
{
"dimension": {
"display_name": "Version Control",
"slug": "version-control"
},
"name": "Git",
"origin": "catalog",
"rationale": "Hygiene: all pipeline and DAG code lives in Git with PR review; universal expectation, weak differentiator (same tiering rationale as both p"
}
]
}
],
"unmapped_skills": []
}
API 2 — extract-details
{}
API 3 — final-role-output
{}
LLM Calls
Every model call made for this run, in pipeline order. Click a card to see the model's response.