After I submitted my thesis at the end of 2024 I looked for some work to do while waiting for the examiners to read it. I ended up at a place called DataAnnotation. They're run by Surge AI, a US-based company that works as a middle-man between several frontier AI labs (and other labs/independent researchers) and a team of specialist contractors that provide training data and human feedback to these labs in various forms.
I work as one of these contractors. Specifically, as a specialist expert in software engineering, mathematics, and physics. The (very large) team of data annotators runs the gamut from generalists whose principal qualification is literacy and an ability to follow instructions, up to certified doctors, lawyers, and PhDs in various fields (me).
I liked the work so much that I kept at it for over a year before even thinking about looking for a "proper job" again.
What I've done
The platform is rather serious about what I am able to disclose. I cannot discuss specifics of most of my projects, though I can give a bit of an overview. My work is broken up into tasks. When I started, these mostly took 1-4 hours, but now I often get tasks that last the full workday or span a whole week or more. Broadly speaking, the tasks are generally to cause, identify, and then closely analyse a failure mode in an LLM. This requires a good understanding of LLMs and their limitations, plus some creativity in coming up with ideas that can challenge them in new ways.
Generally there is a lot of lattitude regarding what we can do with the models. Usually we can pick contexts or codebases that are familiar to us or which we find interesting, though from time to time we work with specific repos or groups of repos. We work with models from most providers, and often I work using bleeding-edge models or agents from frontier labs before they're released to the public, to train them further, provide performance data to the labs, and/or give feedback.
For a few projects, this basically meant I write whatever interested me, provided I did it with these models and it could challenge them. (Thankfully, most of the work I'm interested in challenges LLMs). Generally I cannot share code that has come into contact with these experimental models. I presume this is both to prevent competing labs from understanding the capabilities of new models and to prevent training on these outputs. Sadly this means that I can't share the source code for a few of my personal projects that I've had the chance to work on as part of these tasks.
For other projects, I'm doing software engineering tasks working with large, popular repos to push the models to their limits.
Outside of software, I also did some physics and mathematics work on the platform. The most interesting non-software work I've done is creating difficult benchmark problems that current frontier AIs can't solve, and then validating and solving other people's problems.
What I like
It's really fun! Sometimes it feels like I'm basically doing hobby work, and even for the more structured tasks where I have a little less freedom, it's always interesting and engaging.
I've been able to work with a very wide range of tools and expand my skills. For example, prior to starting this job I had effectively zero webdev experience. I got the chance to do a decent number of tasks working on flask webapps, which gave me the chance to familiarise myself a bit with webdev. (I still don't have particularly much webdev experience, and I'm glad that LLMs are very good at it). For some projects, there was a demand for tasks done in Rust. This was great because I got paid to learn the basics of Rust, and got a bit of experience debugging it when the models failed.
The flexibility was very helpful, particularly at the time I started working. There are no minimum or maximum hours worked per week, so while I was waiting for my thesis to be examined, I worked part-time while doing further research on the side. When it was time for my oral exam, I could drop the work entirely for a period, without having to know in advance how much time I'd need off. Then I could go back up to full time following that. The flexibility was also helpful around the birth of my daughter, as was being able to work odd hours after the kids had gone to bed.
It also pays really well, at least compared to most NZ work.
What I don't like
I'm a contractor, not an employee, so I have basically no rights or benefits. A lot of people complain online that they are just silently kicked off the platform and given no further work without notice, but this has never happened to me because I actually do valuable work. Skill issue I guess.
I don't get things like paid leave (though I can simply just not work when I need to, and the high pay covers it fine). For example of how this has impacted me, recently a requirement was changed which required me (and many other workers) to get a criminal background check done. The third party that did this background check was very good, but I still had to sit around mostly unable to work for about two weeks waiting for the Ministry of Justice to spend 15 seconds confirming that I have no criminal record.
My work is doing something useful, and I appreciate contributing indirectly (and sometimes a little more directly) to the AI tools that I and others use on a near-daily basis. On the other hand, these contributions aren't particularly concrete. I can't really look at any large-scale system and say "this was the part that I built". Going forward, I'd like to do some work where I get a bit more ownership over something I'm building, as I had in my PhD.
In this job, I work alone. It is just me and the model. While I get to review the work of others, this is a very impersonal process. I don't have a fixed manager, or senior mentors who I could learn from. There is a bit of messaging discussion, but it's not the same as true collaboration. I preferred the research environment in this regard, where everyone was an expert in their own domain but worked on overlapping parts of a problem where you could benefit from the expertise of others, and help them using yours.
While remote work has a lot of benefits (and doesn't always have to be as impersonal as this is), actually being physically present in the same place as other people who are good at what they do and who do the same work is something else entirely.