Pipeline run
7cdfa557-83e0-44b7-a607-9ed60d58137b
Client output enrichment
v2 Skill cluster · Nature of work · AI index · Tech stack maturity · Evidence · KRA descriptionvocab breakdown (legacy)
1 LLM mention extraction + deterministic catalog sweep
2 Six-layer resolution cascade (exact → alias → fuzzy → embedding → judge → AI-prediction lane)
3 Table-lookup layering: dim_members → role_dim_map tiers, served verbatim
Data Engineer
in_db name family · data-engineerslug: data-engineer · role_id: 1839a54d-e909-543e-b36e-745551d4e000 · catalog: skill_library_v4 (tiers served verbatim from role_dim_map)
The role is for a Data Engineer, requiring proficiency in Python and SQL for data pipeline development, as well as experience with Apache Spark for data processing. A distinctive requirement is familiarity with FluxNine Orchestrator for scheduling pipelines. The job does not specify a company or domain context.
Job description
Job Title: Data EngineerRequirements:- Python and SQL for pipelines- Apache Spark processing- Experience with FluxNine Orchestrator for pipeline scheduling
Skills from this JD
L0–L3 tiers come verbatim from the role's role_dim_map; violet AI GENERATED chips are AI-prediction-lane skills pending review in the queue.
Apache Beam — Data Processing & Pipeline Frameworks (L0)
Scala — Programming Languages (L0)
Prefect — Workflow Orchestration (L0)
Google Cloud Storage — Cloud Platforms (L1)
Airbyte — Data Ingestion & Integration (L1)
Apache Parquet — Data Lake & Storage Formats (L1)
Library artifacts (this run)
nano JD Parser — gpt-4.1-nano click to toggle
Show raw JSON
{
"JD_type": "pass",
"about_company": null,
"ai_kras": [],
"certifications": [],
"client_details": null,
"company_name": null,
"company_size": null,
"ctc": null,
"domain": {
"primary": {
"aliases": [],
"domain": "Other"
},
"secondary": null
},
"education": [],
"experience": {
"max": null,
"min": null,
"raw": null
},
"job_locations": [],
"notice_period": {
"days": null,
"raw": null
},
"open_to_relocate": false,
"role": "Data Engineer",
"role_aliases": [
{
"name": "Data Engineer",
"reasoning": "generalized form of the picked role",
"relation": "synonym"
}
],
"role_archetype": "Data",
"roles_and_responsibilities": [
{
"bullet_count": 3,
"heading": "Requirements",
"heading_was_present": true,
"source_marker": {
"first_5_words": "- Python and SQL for",
"last_5_words": "for pipeline scheduling"
},
"text": "- Python and SQL for pipelines\n- Apache Spark processing\n- Experience with FluxNine Orchestrator for pipeline scheduling",
"word_count": 19
}
],
"urls": []
}
API 1 — extract-from-jd click to toggle
{
"catalog": "v4",
"consider_adding": [
{
"dimension": "Data Processing \u0026 Pipeline Frameworks",
"reason": "In the role canon\u0027s Data Processing \u0026 Pipeline Frameworks - adding it narrows the candidate pool.",
"skill": "Apache Beam",
"tier": "L0"
},
{
"dimension": "Programming Languages",
"reason": "In the role canon\u0027s Programming Languages - adding it narrows the candidate pool.",
"skill": "Scala",
"tier": "L0"
},
{
"dimension": "Workflow Orchestration",
"reason": "In the role canon\u0027s Workflow Orchestration - adding it narrows the candidate pool.",
"skill": "Prefect",
"tier": "L0"
},
{
"dimension": "Cloud Platforms",
"reason": "In the role canon\u0027s Cloud Platforms - adding it narrows the candidate pool.",
"skill": "Google Cloud Storage",
"tier": "L1"
},
{
"dimension": "Data Ingestion \u0026 Integration",
"reason": "In the role canon\u0027s Data Ingestion \u0026 Integration - adding it narrows the candidate pool.",
"skill": "Airbyte",
"tier": "L1"
},
{
"dimension": "Data Lake \u0026 Storage Formats",
"reason": "In the role canon\u0027s Data Lake \u0026 Storage Formats - adding it narrows the candidate pool.",
"skill": "Apache Parquet",
"tier": "L1"
}
],
"final_skills": [
{
"dimension": {
"display_name": "Programming Languages",
"slug": "programming-languages"
},
"is_primary": true,
"layer": "L0",
"layer_source": "catalog_v4",
"origin": "catalog",
"rationale": "the craft is exercised in Python and SQL; every owned artifact \u2014 transform, DAG, quality check \u2014 is authored in them.",
"skill_id": "646cb350-7c71-5ed2-971e-d141c566ce6e",
"skill_name": "Python"
},
{
"dimension": {
"display_name": "Programming Languages",
"slug": "programming-languages"
},
"is_primary": true,
"layer": "L0",
"layer_source": "catalog_v4",
"origin": "catalog",
"rationale": "the craft is exercised in Python and SQL; every owned artifact \u2014 transform, DAG, quality check \u2014 is authored in them.",
"skill_id": "dd6b38b8-6e50-5e89-a827-5b03688e6113",
"skill_name": "SQL"
},
{
"dimension": {
"display_name": "Data Processing \u0026 Pipeline Frameworks",
"slug": "data-processing-pipeline-frameworks"
},
"is_primary": true,
"layer": "L0",
"layer_source": "catalog_v4",
"origin": "catalog",
"rationale": "transformation engines are where pipeline work physically happens; Spark and dataframe fluency is the single strongest resume signal for the",
"skill_id": "f885334e-9062-5967-bc62-b5b5a27664cf",
"skill_name": "Apache Spark"
},
{
"dimension": {
"display_name": "Workflow Orchestration",
"slug": "workflow-orchestration"
},
"is_primary": true,
"layer": "L1",
"layer_source": "ai_predicted",
"origin": "ai_predicted",
"rationale": "TaBuddy AI prediction \u2014 not in the skill library yet; placed under Workflow Orchestration pending admin review.",
"skill_id": "8d65e0ee-b9f8-5cd9-af6a-b8a3895d4999",
"skill_name": "FluxNine Orchestrator"
}
],
"history_run_id": null,
"jd_parameters": {
"certifications": [],
"clientDetails": null,
"company": null,
"companySize": null,
"ctc": {
"currency": null,
"max": null,
"min": null,
"period": null,
"raw": null
},
"educationRequirements": [],
"experience": {
"max": null,
"min": null,
"raw": null
},
"industryDomain": "Other",
"knockouts": {
"certifications": [],
"ctc": {
"currency": null,
"max": null,
"min": null
},
"educationRequirements": [],
"experience": {
"max": null,
"min": null
},
"location": [],
"noticePeriod": null
},
"locations": [],
"noticePeriod": null,
"openToRelocate": false,
"role": "Data Engineer",
"roleSynonyms": [
"Data Engineer"
]
},
"jd_summary": "The role is for a Data Engineer, requiring proficiency in Python and SQL for data pipeline development, as well as experience with Apache Spark for data processing. A distinctive requirement is familiarity with FluxNine Orchestrator for scheduling pipelines. The job does not specify a company or domain context.",
"layer_conflicts": [],
"nano_parsed": {
"JD_type": "pass",
"about_company": null,
"ai_kras": [],
"certifications": [],
"client_details": null,
"company_name": null,
"company_size": null,
"ctc": null,
"domain": {
"primary": {
"aliases": [],
"domain": "Other"
},
"secondary": null
},
"education": [],
"experience": {
"max": null,
"min": null,
"raw": null
},
"job_locations": [],
"notice_period": {
"days": null,
"raw": null
},
"open_to_relocate": false,
"role": "Data Engineer",
"role_aliases": [
{
"name": "Data Engineer",
"reasoning": "generalized form of the picked role",
"relation": "synonym"
}
],
"role_archetype": "Data",
"roles_and_responsibilities": [
{
"bullet_count": 3,
"heading": "Requirements",
"heading_was_present": true,
"source_marker": {
"first_5_words": "- Python and SQL for",
"last_5_words": "for pipeline scheduling"
},
"text": "- Python and SQL for pipelines\n- Apache Spark processing\n- Experience with FluxNine Orchestrator for pipeline scheduling",
"word_count": 19
}
],
"urls": []
},
"pipeline": "v4",
"rejected": false,
"rejection_code": null,
"rejection_reason": null,
"role": {
"canonical_name": "Data Engineer",
"family": "data-engineer",
"match_method": "name",
"resolution": "in_db",
"role_id": "1839a54d-e909-543e-b36e-745551d4e000",
"similarity": null,
"slug": "data-engineer"
},
"run_id": "jdv4-d541c2b8d0b8",
"secondary_meta": [],
"secondary_skills": [],
"skill_layers": [
{
"label": "Anchor",
"layer": "L0",
"skills": [
{
"dimension": {
"display_name": "Programming Languages",
"slug": "programming-languages"
},
"name": "Python",
"origin": "catalog",
"rationale": "the craft is exercised in Python and SQL; every owned artifact \u2014 transform, DAG, quality check \u2014 is authored in them."
},
{
"dimension": {
"display_name": "Programming Languages",
"slug": "programming-languages"
},
"name": "SQL",
"origin": "catalog",
"rationale": "the craft is exercised in Python and SQL; every owned artifact \u2014 transform, DAG, quality check \u2014 is authored in them."
},
{
"dimension": {
"display_name": "Data Processing \u0026 Pipeline Frameworks",
"slug": "data-processing-pipeline-frameworks"
},
"name": "Apache Spark",
"origin": "catalog",
"rationale": "transformation engines are where pipeline work physically happens; Spark and dataframe fluency is the single strongest resume signal for the"
}
]
},
{
"label": "Primary",
"layer": "L1",
"skills": [
{
"dimension": {
"display_name": "Workflow Orchestration",
"slug": "workflow-orchestration"
},
"name": "FluxNine Orchestrator",
"origin": "ai_predicted",
"rationale": "TaBuddy AI prediction \u2014 not in the skill library yet; placed under Workflow Orchestration pending admin review."
}
]
}
],
"unmapped_skills": []
}
API 2 — extract-details
{}
API 3 — final-role-output
{}
LLM Calls
Every model call made for this run, in pipeline order. Click a card to see the model's response.