Open Source and IP in Tech Due Diligence: The Hidden Legal Risks That Surface After the Deal

Executive summary Most technology due diligence focuses on architecture, code quality, security and the team. Intellectual property often gets less attention, because it sits between the technical and legal workstreams, and neither side fully owns it. Yet IP problems are among the most expensive to discover after closing: a copyleft licence buried in a core module, code written by contractors who never transferred their rights, or a growing share of AI-generated code with unclear ownership. This article explains where these risks typically hide, how to assess them during due diligence, and how to translate findings into deal terms rather than post-acquisition surprises.
Why IP Risk Deserves Its Own Workstream
In a typical transaction, the legal team reviews contracts, trademarks and patents, while the technical team reviews the codebase. The gap between the two is exactly where software IP risk lives. Lawyers rarely scan dependency trees, and engineers rarely read employment agreements or contractor terms.
What makes IP findings different from most technical issues is their timing and impact. Technical debt slows a roadmap; an IP defect can undermine the very asset you paid for. Common consequences include:
- a forced release of proprietary source code under a copyleft licence,
- costly re-engineering to replace a component the company is not allowed to use commercially,
- disputes with former contractors or founders over who owns the code,
- breaches of reps and warranties that trigger claims long after the purchase price has been paid.
These problems rarely appear in the first weeks after closing. They usually surface when the buyer tries to do something new: sell to an enterprise customer with strict procurement, prepare for the next exit, or integrate the product into a larger platform.
Not All Open Source Licences Are Equal
Almost every modern software product is built largely on open source. That is not a problem in itself. The risk depends on which licences are used, how the code is used, and how the product is distributed.
Two points are frequently misunderstood.
First, the "SaaS loophole" does not cover everything. Many teams assume that because they do not distribute software, GPL obligations never apply. That holds for GPL in a pure SaaS model, but not for AGPL, and not for components shipped to customers as on-premise installers, mobile apps, desktop clients, SDKs or embedded firmware.
Second, licences change over time. Several widely used infrastructure projects have moved from open source to more restrictive source-available licences in recent years, HashiCorp's shift to the Business Source License in 2023 being a prominent example. A dependency that was harmless when adopted may carry very different terms in its current version.
Where the Risks Actually Hide
A clean list of top-level dependencies does not mean a clean codebase. In our experience, the most material findings tend to come from less obvious places.
Transitive dependencies. A permissively licensed library can pull in dependencies with far more restrictive terms. Unless the full dependency tree is analysed, these stay invisible.
Copied snippets. Code pasted from public repositories, forums or tutorials does not appear in a package manifest. Detecting it requires snippet-level scanning, not just manifest analysis.
Vendored and forked code. Libraries copied directly into the repository, often modified years ago, lose their visible link to the original project and its licence.
Commercial components. Paid libraries, SDKs, fonts and datasets come with their own licence terms. Some are tied to a specific legal entity, number of users or deployment model, and may not survive a change of control without renegotiation.
Missing notices. Even permissive licences require attribution. Missing notice files are usually easy to fix, but they signal how mature the company's overall compliance process is.
Who Actually Owns the Code?
Open source is only half of the picture. The other half is whether the target company truly owns the code it wrote itself. Ownership gaps are common in fast-growing companies that relied on freelancers, agencies or founders working before the company was incorporated.
Key questions to answer include:
- Did every employee and contractor sign an agreement that effectively transfers intellectual property rights to the company?
- Was code written by founders before incorporation formally assigned to the company?
- Do agreements with software houses and agencies transfer ownership, or merely grant a licence?
- Are there university, grant or accelerator agreements that give third parties rights to parts of the technology?
Jurisdiction matters here. In the United States, work made for hire by employees generally belongs to the employer. In many European countries, the rules are stricter, especially for contractors. In Poland, for example, copyright in software created by a B2B contractor does not transfer automatically. A valid transfer requires a written agreement that specifies the fields of exploitation. As many DACH and international companies build teams through Polish B2B contracts, this is one of the most frequent and most easily overlooked findings in cross-border deals.
AI-Generated Code: The New Category of IP Risk
In 2026, it is rare to find an engineering team that does not use AI coding assistants. This creates questions that did not exist in due diligence checklists a few years ago.
Copyright protection may be limited. In many jurisdictions, including the United States and the EU member states, copyright protection requires human authorship. The US Copyright Office has confirmed that purely AI-generated material is not protected, while human contributions such as selection, arrangement and meaningful modification can be. For most products this is not a crisis, since engineers review, edit and integrate AI output. The risk grows where large, self-contained modules were generated with minimal human input.
Suggestions can reproduce licensed code. AI tools occasionally output code that closely matches their training data, including code under copyleft licences. Many enterprise tools offer filters that block suggestions matching public code, but only if they are switched on.
Tool terms and data handling vary. Consumer-grade tools may have different terms regarding ownership of output and use of submitted code than enterprise plans. If developers pasted proprietary code or customer data into unapproved tools, this may also raise confidentiality and data protection concerns.
When assessing a target, the question is not whether the team uses AI, but whether it uses AI in a governed way. Look for:
- an approved list of AI tools with enterprise terms,
- enabled duplicate-detection or public-code filters,
- a written policy on what code and data may be shared with AI tools,
- review standards that apply equally to AI-assisted and manually written code.
A team that answers these questions confidently usually has a mature engineering culture overall. A team that cannot answer them often has similar gaps elsewhere. We describe how to put these rules into practice on our page on AI-assisted software delivery.
How to Assess IP Risk During Due Diligence
An effective IP assessment combines automated analysis with document review and targeted interviews. A practical approach looks like this.
1. Generate a software bill of materials (SBOM). Use software composition analysis (SCA) tools to build a complete inventory of direct and transitive dependencies, including licence information. Standard formats such as SPDX or CycloneDX make the results reusable. This will soon matter beyond M&A: the EU Cyber Resilience Act introduces SBOM-related obligations for products with digital elements, with most requirements applying from December 2027. Companies that build licence scanning and SBOM generation into their CI/CD pipelines and cloud operations are far better prepared for both.
2. Run snippet-level scanning on critical repositories. Focus on the core product rather than every internal tool. The goal is to find copied code without licence attribution.
3. Map licences against the distribution model. The same licence can be harmless in a backend service and highly problematic in a mobile app or on-premise installation. Findings must be assessed in context.
4. Review ownership documentation. Cross-check commit history against the list of employees and contractors with signed IP agreements. Contributors who appear in the codebase but not in the contract records are an immediate red flag.
5. Review commercial licences for change-of-control clauses. Identify components whose licences terminate or require consent when the company is acquired.
6. Assess AI usage and governance. Interview engineering leads, review tool configurations and check whether an AI usage policy exists and is followed in practice.
Turning Findings into Deal Terms
Not every finding is a deal breaker. The goal of IP due diligence is not a perfect codebase, but an accurate price and appropriate protection. Findings typically fall into three groups.
Low severity: fix after closing. Missing attribution notices, outdated licence files and minor permissive-licence gaps. These belong in the 100-day plan.
Medium severity: fix before closing or price accordingly. Weak copyleft components used in a way that may create obligations, commercial licences requiring renegotiation, or missing IP assignments from a small number of former contractors. Typical instruments include pre-closing conditions, specific indemnities and remediation cost estimates reflected in the valuation.
High severity: fundamental to the investment thesis. Strong or network copyleft code embedded in core proprietary modules of a distributed product, or significant parts of the codebase with unclear ownership. These may justify a meaningful price adjustment, escrow arrangements, extended warranties or, in extreme cases, walking away.
The key is to quantify remediation: how many engineering weeks would it take to replace a problematic component, re-implement a module or obtain missing assignments? That estimate turns an abstract legal risk into a number both sides can negotiate. Where remediation means reworking larger parts of the product, an AI refactoring assessment helps size the effort before signing, and the work itself can be planned as part of ongoing product and application engineering without stalling the roadmap.
Clean IP Is Part of the Asset You Are Buying
Buyers pay for software because they expect to own it, use it and build on it freely. If ownership is unclear or licences restrict future plans, part of that value simply does not exist. The good news is that most IP issues are detectable before signing, and many are inexpensive to fix when found early.
The most effective approach treats IP as a joint technical and legal workstream, supported by automated scanning, contract review and an honest look at how the team uses AI. For sellers, the same exercise is one of the simplest ways to protect valuation before a buyer's due diligence begins. If you are preparing a transaction or an exit, get in touch to see how our technology due diligence works in practice.
This article is for general information and does not constitute legal advice. Licence and ownership questions should always be reviewed with qualified legal counsel in the relevant jurisdiction.
FAQ - Open Source and IP in Tech Due Diligence
Does using GPL-licensed software automatically make a SaaS product open source?
No. The GPL is triggered mainly by distribution, so software running only on your servers is generally not affected. The AGPL is different: it can require source disclosure when users interact with the software over a network. Any components shipped to customers, such as mobile apps or on-premise installers, must also be reviewed separately.
How long does an IP and open source review take during due diligence?
For a typical mid-sized SaaS product, automated scanning and an initial licence report can be completed within a few days. Ownership review and interviews usually run in parallel with the broader technical due diligence, so the full assessment typically fits within the standard DD timeline.
Is AI-generated code a reason to lower the valuation?
Not on its own. What matters is governance: approved tools, enabled public-code filters, clear usage policies and consistent human review. Uncontrolled use of consumer AI tools on core code is a risk signal; structured use of enterprise tools is increasingly a sign of engineering maturity.
What is the most common IP finding in cross-border deals involving Central European teams?
Missing or incomplete IP transfer agreements with contractors. In countries such as Poland, rights in software created under a B2B contract do not pass to the client automatically, so agreements must explicitly transfer them and list the relevant fields of exploitation.
Can IP issues be fixed after the transaction closes?
Many can. Missing notices, most permissive-licence gaps and some ownership issues are straightforward to resolve. The problem arises with high-severity findings, such as copyleft code in core modules, which may require substantial re-engineering. These should be identified and priced before signing, not discovered during integration.
Should sellers run their own IP review before going to market?
Yes. A pre-exit review lets the company fix issues on its own terms, prepare clear documentation for the data room and avoid findings that a buyer could use to renegotiate the price late in the process.



