AI Shame in Software Teams: Why Engineers Hide AI Use
Engineers hide AI use because teams still price visible effort and penalize disclosure. The ethical lines are accountability, data boundaries, and honest evaluation.
Most software engineers use an AI assistant, and a large share of them keep quiet about it. Slack’s Fall 2024 Workforce Index found that 48% of desk workers would be uncomfortable telling their manager they used AI for common tasks, including writing code. The DORA 2025 report puts AI adoption among software professionals at 90%. The two surveys measure different populations, but the gap between them has a cause. My position: hiding AI use is a rational response to a real social penalty, so the fix is a team norm, not individual guilt. Only three lines deserve judgment: accountability for shipped code, data staying inside approved tools, and honest representation when a skill is being evaluated. Everything else is tooling.
Why engineers keep AI use to themselves#
The Slack survey asked the uncomfortable people why. The top reasons were that it feels like cheating (47%), that it might make them look less competent (46%), and that it might make them look lazy (46%). Company policy came last at 21%. In practice, the barrier is social long before it is regulatory. Microsoft’s 2024 Work Trend Index, run with LinkedIn across 31,000 people in 31 countries, points the same way. Among people who use AI at work, 52% are reluctant to admit using it for their most important tasks, and 53% worry that it makes them look replaceable.
Ethan Mollick gave the pattern a name in 2023: “secret cyborgs”. His informal poll suggested that over half of generative AI users had used it without telling anyone, at least some of the time. He listed three reasons: outright bans, the suspicion that AI-written work is judged by a different standard, and the fear that visible productivity gains invite layoffs. None of the three is about the quality of the work.
The root is older than the tools. Many teams still evaluate engineers on visible effort: hours in the editor, lines in the diff, the struggle that was witnessed. An assistant breaks the link between effort and output. While that scoreboard is still running, an engineer who finishes a task in an afternoon with help can look slow or look like a cheat. Silence is the obvious third option.
The penalty in controlled experiments#
The fear is not imaginary. Reif, Larrick, and Soll published four preregistered experiments in PNAS in 2025, with 4,439 participants in total. People who received help from AI were rated lazier, less competent, and less diligent than people who received help from a human or whose source of help was unstated. The effect held across age, gender, and occupation, and it reached hiring judgments. Managers who rarely used AI themselves were less willing to hire a candidate described as using AI daily.
Two boundary conditions in the same study matter more for leads than the headline. The penalty was no longer significant among evaluators who used AI weekly or daily. And managers who used AI frequently favored candidates who also used it. The judgment is therefore a property of the evaluator, which is something a team can change.
Disclosure has its own cost. Schilke and Reimann ran 13 experiments, published in Organizational Behavior and Human Decision Processes in 2025. People who disclose AI use are trusted less than people who do not, because the work is seen as less legitimate. Voluntary or mandatory made no difference, and being exposed by a third party was worse than disclosing yourself. So the advice to be transparent is honest, with a price attached. That price lands on the individual unless the team absorbs it. The same result makes hiding a poor long-term bet: in a team with code review and git history, exposure is likely, and exposure costs more.
The strongest case against AI-assisted code#
The objection deserves its best form before any answer. Software is a comprehension discipline. Ghostty’s AI policy states the requirement plainly: “The human-in-the-loop must fully understand all code.” If an author cannot explain a change, the reviewer becomes the only person who understands it, and often not even them. Self-assessment is unreliable here. METR’s randomized trial from early 2025 gave 16 experienced open-source developers 246 tasks in their own repositories. With AI allowed, they were 19% slower, while believing afterwards that they had been 20% faster. People who misjudge their own speed by that margin may misjudge their own understanding too. The cost also moves. Stack Overflow’s 2025 survey found the top frustration with AI tools is solutions that are almost right but not quite (66%), and almost-right code is cheap to produce and expensive to review. Finally, there is the deskilling concern for junior engineers: skills that are never exercised do not grow, and an assistant removes the exercise.
Every point in that list is correct, and every point is an argument for accountability and for disclosure. None of them is an argument for shame. Shame pushes use underground, which removes exactly the signal a careful reviewer needs: which parts were generated, which were rewritten, and where the author’s own understanding is thinnest. The comprehension objection is answered by a review question, such as asking the author to walk through the retry path. That question works only when the reviewer knows where to ask it. Hidden use hides that gap; the judgment an assistant cannot supply is the subject of Phronesis and AI coding agents.
The speed number itself deserves care. METR’s February 2026 update reports that more developers now decline to join the study because they do not want to work without AI, which weakens the newer data. METR itself believes developers are now likely sped up by AI tools, and still rates its evidence on the size of any speedup as weak. The durable finding is the gap between felt speed and measured speed.
Three lines that belong in work ethics#
Is it legitimate to judge someone for using AI? Judge the work against three lines. Nobody discloses autocomplete, a Stack Overflow answer, or an IDE refactor, because origin says nothing a reviewer needs. The useful question for a reviewer or an evaluator is who owns the code now. That framing is the same one used for inherited code in accountability versus blame: authorship is history, ownership is present tense.
The first line is accountability. The ACM/IEEE-CS Software Engineering Code puts the duty to accept full responsibility for one’s own work in its first clause (1.01), and the ACM Code of Ethics (2.1) asks for high quality in both process and product. The Linux kernel’s guidance on coding assistants turns the principle into procedure. An AI agent must not add a Signed-off-by line, and the human submitter reviews all generated code and takes full responsibility for the contribution. Shipping code you cannot explain or defend is a professional failing. Using a tool to draft it is not.
The second line is the data boundary. ACM Code 1.7 asks members to honor confidentiality, and 2.3 asks them to know and respect the rules that apply to professional work. Pasting proprietary source into a public chatbot breaks both. Samsung found this in 2023, when engineers did exactly that and the company restricted generative AI in response. Microsoft’s Work Trend Index found that 78% of AI users bring their own tools to work, which says where this line fails in practice: when a team offers no approved tool, the personal account becomes the default. The governance side is covered in AI coding tools: security risks and governance.
The third line is honest representation of what is being evaluated. ACM Code 1.3 asks members to be honest and trustworthy, including about their own qualifications. In concrete terms: if an interview, a take-home, or a certification measures unaided skill, using AI silently is misrepresentation. If the evaluation measures outcomes or collaboration with AI, it is not. Both rules can be legitimate. The condition is that each is stated.
Everything outside the three lines (drafting, boilerplate, test scaffolding, explaining an unfamiliar API) is tooling, and nobody’s business by default.
What a software team can do#
A published stance, modeled from the top#
DORA’s AI Capabilities Model lists a clear and communicated AI stance first among its seven capabilities. DORA reports that the stance amplifies AI’s positive effects and can reduce friction for employees, whatever the policy says. Slack’s data points the same way: workers who felt comfortable sharing their AI use were 67% more likely to have used AI for work. A page that names the approved tools, says what data may go where, and states that AI assistance carries no evaluation consequence removes most of the guesswork that drives hiding.
The stance needs a visible owner. The penalty in Reif’s experiments was no longer significant among evaluators who used AI weekly or daily. The cheapest intervention, then, is a lead saying in a visible channel that an assistant drafted this RFC, and here is what changed. That sentence lowers the cost for everyone below. The trade-off is partial exposure: leads tend to disclose AI-polished prose and never AI-generated code, so the signal stays half-sent. The disclosure has to include the kind of work engineers are hiding. That sentence is a communication act of the kind described in communication as the gate to the next level.
Disclosure as review metadata#
The Linux kernel’s convention is a working model for product teams. An Assisted-by trailer records AI assistance in one line, and responsibility is unchanged. For a product team, the equivalent is a short free-text field in the pull request template, labelled AI assistance. It asks which tool, which parts of the diff, and what the author changed or rejected. Leaving it empty is allowed. Its only purpose is to direct review attention, and the template should say so in the field’s own help text.
A generalized example shows why. An engineer drafts a data migration with an assistant, rewrites about half of it, and opens the PR with a description that mentions neither. The reviewer assumes the retry path was reasoned through, approves after a light pass, and a condition the assistant invented reaches staging. One line naming the generated section would have put the reviewer’s attention where it was needed. The review went worse because the metadata was missing. The author’s silence was a rational reaction to a team that had never said disclosure was safe.
The failure mode of the field is that it becomes a blame filter. If AI-assisted PRs get extra scrutiny and the rest sail through, people stop filling it in, and the team is back to hidden use. The mitigation is explicit: the field steers review focus and never appears in evaluation. Ghostty sits at the other end: disclosure required from outside contributors, maintainers exempt, and a public denouncement list of contributors who submit AI slop. Disclosure as a gate is defensible under open-source spam pressure and would read as distrust inside a product team.
Hiring rules stated per stage#
Canva insists that candidates use AI in its coding interviews and evaluates how they collaborate with it. Anthropic’s candidate guidance asks for no AI in take-homes and live interviews unless a stage says otherwise, and asks candidates to be transparent. Two AI-forward companies, two opposite rules, and both are honest for the same reason: each states what the stage measures. The trade-off is that honest rules require interview redesign. Canva rebuilt its format. A rule stated once on a careers page and never repeated in the interview itself does not count as stated.
Performance reviews without an AI question#
Shopify’s memo makes “reflexive AI usage” a baseline expectation and adds AI usage questions to performance and peer reviews. The intent is clear, and it does end hiding quickly. My recommendation is against it. Grading the act of use produces performative use. Engineers who judge that an assistant is wrong for a task feel pressure to use it anyway, and the review question ends up measuring compliance with a tool choice. The review should grade outcomes, judgment, and the ability to explain and defend a change, with the three lines as the only hard boundaries. The trade-off is calibration. Outcomes are harder to grade than effort proxies. The quiet failure is outcomes sliding into output volume, which an assistant inflates for free. The guard is the explain-and-defend criterion, because volume without understanding fails it.
When the default does not hold#
The default of cheap, consequence-free disclosure plus three explicit lines assumes a product team with near-universal AI use and routine code review. It also assumes the team avoids the small habits that rebuild the penalty: treating the PR field as a confession, policing where code came from when the useful question is whether the author can explain it, and reading self-reported speedups as proof. Three situations override the default. Assessments that explicitly measure unaided skill (interviews, take-homes, some certifications) call for abstaining or for disclosing before submission. Client or regulatory contracts that restrict tool use or require disclosure take precedence over team habit. Open-source projects with their own policy, like the kernel or Ghostty, get their rule followed whatever the contributor’s home team does. Outside those three, whether someone used AI stays review metadata, and judgment is reserved for what shipped, where the data went, and what was claimed. A reasonable first step for a lead is to put the one-line disclosure on their own next pull request and watch what the team does after that.
References#
- Evidence of a social evaluation penalty for using AI (PNAS, 2025) (opens in new tab) - Reif, Larrick, and Soll’s four preregistered experiments (N = 4,439): AI users are judged lazier and less competent, and the effect reaches hiring decisions.
- Is AI Damaging Your Professional Image? (Duke Fuqua Insights) (opens in new tab) - Plain-language summary of the PNAS study, including the finding that the penalty disappears when the evaluator uses AI.
- The transparency dilemma: How AI disclosure erodes trust (OBHDP, 2025) (opens in new tab) - Schilke and Reimann’s 13 experiments: disclosure lowers trust through perceived legitimacy, and third-party exposure is worse than self-disclosure.
- Detecting the Secret Cyborgs (One Useful Thing) (opens in new tab) - Ethan Mollick’s essay naming workers who use AI without telling anyone, with the reasons and what organizations can do.
- The Fall 2024 Workforce Index (Slack) (opens in new tab) - Survey of 17,372 desk workers: 48% uncomfortable admitting AI use to their manager, the reasons behind it, and the 67% comfort-to-use link.
- AI at Work Is Here. Now Comes the Hard Part (Microsoft Work Trend Index 2024) (opens in new tab) - Microsoft and LinkedIn report on bring-your-own-AI (78%) and reluctance to admit AI use for important tasks (52%).
- 2025 Stack Overflow Developer Survey: AI (opens in new tab) - Developer adoption (84%), daily use among professionals (51%), and the “almost right” frustration (66%).
- How are developers using AI? Inside our 2025 DORA report (Google) (opens in new tab) - 90% AI adoption among software professionals and the split in trust that comes with it.
- Introducing DORA’s inaugural AI Capabilities Model (Google Cloud) (opens in new tab) - Seven capabilities, led by a clear and communicated AI stance that reduces friction regardless of policy content.
- Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity (METR) (opens in new tab) - Randomized trial where developers were 19% slower with AI while believing they were 20% faster.
- We are Changing our Developer Productivity Experiment Design (METR) (opens in new tab) - Follow-up explaining why newer data is weaker, including developers declining to work without AI.
- AI Coding Assistants (Linux kernel documentation) (opens in new tab) - Kernel rules: no Signed-off-by from an agent, full human responsibility, and the Assisted-by tag.
- Ghostty AI Policy (opens in new tab) - Open-source policy requiring disclosure of AI use by outside contributors and full human understanding of submitted code.
- Yes, You Can Use AI in Our Interviews. In fact, we insist (Canva Engineering) (opens in new tab) - Why Canva requires AI in coding interviews and expects full ownership of the resulting code.
- How to collaborate with Claude during our hiring process (Anthropic) (opens in new tab) - Stage-by-stage rules: AI for preparation and refinement, no AI in take-homes or live interviews unless stated.
- Reflexive AI usage (Shopify) (opens in new tab) - Tobi Lütke’s memo making AI use a baseline expectation and adding it to performance and peer reviews.
- ACM Code of Ethics and Professional Conduct (opens in new tab) - Principles on honesty (1.3), confidentiality (1.7), quality of work (2.1), and respecting rules for professional work (2.3).
- Software Engineering Code of Ethics and Professional Practice (ACM/IEEE-CS) (opens in new tab) - Joint code whose first clause asks engineers to accept full responsibility for their own work.
- Samsung puts ChatGPT back in the box after ‘code leak’ (The Register) (opens in new tab) - Case where engineers pasted proprietary code into a public chatbot, illustrating the data-boundary line.
Related posts
A field guide to spotting, managing, and resolving conflict in software teams, with practical frameworks and early-warning systems that turn friction into performance.
leadership · team-management · best-practices +4
Stop asking who wrote the legacy code. Separate responsibility, accountability, and blame, and make inherited code owned rather than orphaned.
engineering-culture · leadership · team-management +3
A blameless postmortem model that fixes the system instead of finding a culprit, with a copy-paste template and where individual accountability still applies.
engineering-culture · incident-response · psychological-safety +4
Why the salary-vs-impact-vs-satisfaction debate is a false trichotomy, and what your answer reveals about your vocational development stage.
career · team-dynamics · developer-experience +1
A field guide to engineering-specific difficult coworkers, from code-review blockers to ghost colleagues, with practical strategies that work for each archetype.
leadership · team-management · best-practices +5