Obiguard logo

AI Safety Researcher

Part-time

Obiguard

Remote · MYFull-time / Part-time / InternshipPosted 3w agoStill listed 2 days ago

Most applications go out cold — see where you stand first. No sign-up to start.

Watch jobs like this.

At a glance

Compensation
No compensation found
Location
Remote · MY
Schedule
Full-time / Part-time / Internship
Work Authorization
Not specified

Olive lists jobs from US employers, including remote roles you can work from the United States.

Job overview

The AI Safety Researcher will monitor emerging attacks on large language models and autonomous agents, reproduce these attacks, and collaborate with engineering to develop detectors that are rapidly deployed. The role includes on‑site customer pilots for up to three weeks, publishing findings, and mapping attacks to compliance frameworks such as NIST AI RMF and the EU AI Act.

Skills & qualifications

RequiredNice to have

Skills

ML/NLPApplied Security ResearchAdversarial Machine LearningPromptingFine‑TuningRed‑Teaming LLMsPython ProgrammingEvaluation Tooling DevelopmentTechnical CommunicationPublicationsCTFsBug BountiesLangGraphAutoGenMCP‑Based Tooling

Full job description

Department RESEARCH

Location Kuala Lumpur · Remote-friendly

Type Full-time Open to part-time arrangements, and we welcome internship applications for this role.

How to apply Email your CV to [email protected] with the role title in the subject line. Apply now →

About the role Attacks against LLMs and autonomous agents evolve weekly — new jailbreaks, new injection techniques, new exfiltration patterns. As AI Safety Researcher, you’ll track this landscape ahead of our customers, reproduce attacks against real model deployments, and work with engineering to turn findings into detectors that ship into the inspection pipeline within days, not quarters. This is a forward-deployed research role — you may be deployed on-site with customers during pilots and security reviews, for stretches of up to three weeks at a time, advising on emerging risks in person.

What you'll do

  • Research and reproduce emerging attack techniques against LLMs and autonomous agents — prompt injection, jailbreaks, data exfiltration, tool-call abuse.
  • Design evaluation suites and red-team harnesses to continuously test Obiguard’s own detectors against novel attacks.
  • Translate research findings into production detector specs and policy primitives, working closely with the engineering team.
  • Publish internal and external write-ups on notable findings, contributing to Obiguard’s standing as a research-driven vendor.
  • Map new attack classes and detectors to relevant compliance frameworks (NIST AI RMF, ISO 42001, EU AI Act).
  • Advise customers on emerging risks during pilots and security reviews.

What we're looking for

  • Strong background in ML/NLP, applied security research, or adversarial machine learning — academic or industry.
  • Hands-on experience prompting, fine-tuning, or red-teaming LLMs.
  • Comfortable reading and reproducing findings from security research papers and disclosures.
  • Able to communicate technical findings clearly to both engineers and non-technical stakeholders.
  • Programming proficiency in Python; comfortable building quick evaluation tooling from scratch.
  • Comfortable with extended client-site deployments — this is a forward-deployed role, with on-site stints of up to three weeks at a time during active engagements.

Nice to have

  • Publications or public write-ups on LLM/agent security, adversarial ML, or red-teaming.
  • Experience with CTFs, bug bounties, or formal security research.
  • Familiarity with agent frameworks (LangGraph, AutoGen, MCP-based tooling).

Similar jobs, posted recently

Open roles like this one, listed in the last 30 days.

You've read the whole posting — now see how you match it.