Proof of Skill logo

Graph RAG - Knowledge Engineer

Proof of Skill

Hyderabad, Telangana, IndiaJobPosted 3 days agoStill listed 3 days ago

Most applications go out cold — see where you stand first. No sign-up to start.

Watch jobs like this.

At a glance

Compensation
No compensation found
Location
Hyderabad, Telangana, India
Work Authorization
Not specified

Olive lists jobs from US employers, including remote roles you can work from the United States.

Job overview

Proof of Skill’s Chryselys team seeks a Knowledge Engineer to design, build, and operate Python/FastAPI services that extract entities and relationships from unstructured pharma documents, resolve them to canonical identifiers, maintain a large‑scale knowledge graph, and integrate graph‑augmented retrieval with vector search for multi‑hop queries.

Skills & qualifications

RequiredNice to have

Skills

Knowledge Graph ConstructionNatural Language ProcessingEntity ExtractionEntity ResolutionGraph Database (Neo4j)Amazon NeptuneCypherOpenCypherOntology ModelingTaxonomy ModelingNamed Entity RecognitionRelation ExtractionLarge Language ModelsspaCyscispaCyTransformersPythonFastAPIPydanticDockerPytestGitCIAWSS3ECS/EKSBedrockBiomedical Ontologies

Qualifications

5–9 Years Software Engineering Experience

Full job description

About Chryselys

Chryselys is a Great Place to Work Certified Pharma Analytics & Business consulting company that delivers data-driven insights leveraging AI-powered, cloud-native platforms to achieve high-impact transformations. We specialize in digital technologies and advanced data science techniques that provide strategic and operational insights. Role SummaryDesign, build and operate the Python/FastAPI services that extract entities and relationships from unstructured documents, resolve them to canonical identifiers, maintain the knowledge graph, and serve graph-augmented retrieval alongside vector search for multi-hop and relational questions. Responsibilities • Build entity and relation extraction services over unstructured documents — molecules, brands, indications, therapeutic areas, endpoints, claims.

  • Build entity resolution: alias handling, blocking and candidate generation, fuzzy and embedding matching, calibrated thresholds, human review routing.
  • Design and maintain the graph schema and ontology; incremental ingest, node and edge deduplication and merging, provenance on every edge.
  • Fuse graph and vector results into a single ranked, cited context for the retrieval service.
  • Instrument, monitor and support the services in production.

Qualifications • 5–9 years software engineering, with demonstrable knowledge-graph construction and applied NLP delivered to production.

  • Has built a knowledge graph from unstructured text — not queried an existing one, and not a CRUD application on a graph database.
  • Graph at production scale. Millions of nodes and edges; incremental updates with stable node identity; supernode and traversal-explosion handling with bounded depth and timeouts.
  • Entity resolution at corpus scale. Blocking and candidate generation that avoid O(n²) comparison, with measured precision on a labelled sample.
  • Graph database in production. Neo4j, Amazon Neptune or equivalent; Cypher / openCypher fluency.
  • Ontology and taxonomy modelling. Schema evolution without breaking downstream consumers; judgement on node vs. edge vs. property.
  • Extraction. NER and relation extraction — LLM-based, model-based (spaCy, scispaCy, transformers) or hybrid, with the judgement to choose.
  • Graph vs. vector judgement. Knows where graph retrieval wins — multi-hop, relational, comparative and aggregate questions — and that hybrid is the production norm.
  • Python and FastAPI in production. Python 3.11+, async, Pydantic, Docker, pytest, Git and CI; AWS as a consumer (S3, ECS/EKS, Bedrock, Neptune or self-hosted Neo4j).

Preferred • Biomedical ontologies and registries: UMLS, MeSH, SNOMED, RxNorm, ICD-10, DrugBank, ChEMBL.

  • Life sciences or pharma domain experience; RDF/SPARQL alongside property graphs.
  • GraphRAG approaches: community detection for corpus-level summarisation, local vs. global search.

Similar jobs, posted recently

Open roles like this one, listed in the last 30 days.

You've read the whole posting — now see how you match it.