AI companies are pushing the boundaries of data collection, and it's sparking a controversial debate! OpenAI is requesting contractors to upload their past work assignments, a bold move to enhance AI agent performance. But is it a step too far?
According to WIRED's investigation, OpenAI is seeking real-world tasks and assignments from third-party contractors to evaluate its AI models' capabilities. This initiative is part of OpenAI's ambitious plan to create a human baseline for various tasks, which can then be compared to AI performance. In September, they introduced a new evaluation process to measure AI models against human experts in different industries, claiming it as a milestone towards achieving AGI (Artificial General Intelligence).
Here's the twist: OpenAI wants contractors to provide actual work examples, not just hypothetical scenarios. They are instructed to upload concrete outputs, such as Word documents, PDFs, or even code repositories, from their current or previous jobs. And this is where it gets controversial—contractors are also allowed to share fabricated work, as long as it realistically represents their skills.
The project has a two-fold structure: task requests and task deliverables. OpenAI emphasizes the need for 'real, on-the-job work,' ensuring that the examples are authentic. For instance, a task for a luxury concierge company involves creating a yacht trip itinerary, and the deliverable is a real itinerary created for a client.
But there's a catch. OpenAI asks contractors to remove sensitive information, including personal data and corporate secrets. This has raised concerns about potential trade secret misappropriation and NDA violations. Intellectual property lawyer Evan Brown warns that AI labs must be cautious, as they could be held liable if confidential information is not properly protected.
AI labs are investing heavily in this data-driven approach, with companies like OpenAI, Anthropic, and Google hiring large numbers of contractors to generate high-quality training data. The demand for skilled talent has created a booming sub-industry, with companies like Handshake AI and Surge valued at billions of dollars.
OpenAI has also explored other avenues for obtaining real company data, including reaching out to businesses selling assets after going out of business. However, concerns about personal information protection have hindered these efforts.
And this is the part most people miss: the fine line between AI advancement and ethical data handling. As AI companies strive for more realistic training data, how can they ensure they're not crossing legal and ethical boundaries? It's a complex issue that demands attention, and we want to hear your thoughts. Do you think AI labs are taking enough precautions to protect sensitive information? Are they striking the right balance between innovation and privacy? Share your opinions below, and let's spark a thoughtful discussion!