A featured contribution from Leadership Perspectives: a curated forum reserved for leaders nominated by our subscribers and vetted by the CIOReview Advisory Board.

Commonwealth of Virginia

Who Owns Government Data?

It seems straightforward and simple that taxpayers own government data, but in practice it’s a bit more complex. Public sector organizations collect, process, manage, analyze, interpret, utilize, and publish all types of data in support of their day-to-day operations. How data is collected, its intended use, and the attributes contained within it are some of the criteria used to describe its sensitivity, security classification, and, in the Commonwealth of Virginia, its data tier. Sensitivity is primarily focused on whether the data asset contains confidential information (PII, PHI, PCI, etc.), whether the integrity of the data asset could be compromised, and whether the dataset is available to support critical mission processes or CIA (sensitive with regards to Confidentiality, Integrity, or Accessibility).

From these attributes, Virginia categorizes data assets into 5 tiers from 0 to 4 as defined in the Commonwealth of Virginia Data Trust. Tier 0 data is neither sensitive nor proprietary and intended for public access. This data is well suited for publishing on Virginia’s Open Data Portal. Tier 1 data is similar to Tier 0 in that it is not protected from public disclosure or subject to withholding under any law, regulation, or contract, however, publication of Tier 1 data on the public Internet would have the potential to jeopardize the safety, privacy, or security of a person who may be identified through the use of said data. This data tier is usually made available through FOIA requests and is often redacted or suppressed to limit the ability of the recipient to identify individuals or population subgroups by combining with other publicly available data assets (otherwise known as the mosaic effect).

"Technology procurement contract language should include provisions for the public sector entity to have direct access to all data stored and managed by the vendor either through APIs or some other programmatic means."

Tier 2 data includes sensitive and proprietary information intended for access or release only on a ‘need-to-know’ basis, including personal information not otherwise classified as Tier 1 or 0, and data protected or restricted by contract, grant, or other agreement terms and conditions. These are data assets and elements that are critical to the mission of the organization in service delivery to constituents. Examples include vaccine dose distributions to providers, vaccine administrations by provider, de-identified COVID case and vaccine administrations, foster care demographics, de-identified EMS substance use incidents, and others. Although not technically PII, data in this tier comes pretty close.

Tier 3 data assets contain sensitive or proprietary information and data elements with a statutory requirement under relevant

state and federal laws to notify affected parties in case of a confidentiality breach. This data includes PII, PHI, PCI, FERPA and other protected information classes including social security numbers, driver’s license numbers, financial account information, personal medical information, etc. This is the tier we usually think about when discussing protected data, however, there are other examples which include attorney-client privileged information, criminal justice information, critical infrastructure, and protected tax information that also fall within this tier.

Lastly, there’s Tier 4 data which includes sensitive or proprietary data where the unauthorized disclosure could potentially cause major damage or injury, including death, to entities or individuals identified in the information or otherwise significantly impair the ability of the public sector organization to perform its statutory functions. By default, any dataset designated by a federal agency at the “Confidential” level or higher under the federal government’s system for marking classified information is considered Tier 4. The FBI’s Criminal Justice Information Services (CJIS) is a good example of tier 4 data.

With 5 data tiers classifying data in hundreds of information systems supporting thousands of business processes, it is easy to see how establishing ‘ownership’ can be a complex endeavor. As mentioned earlier, ultimately the taxpayers own government data, but they have entrusted the public sector organizations as stewards of data to properly deliver government services. In the context of data governance, public sector organizations define data owners as those individuals that have control of the information systems that collect, house, and manage the data in question and make decisions about who can access that data and for what purpose. Public sector data owners oversee the development and operation of data systems, determine access privileges, and participate in governance processes.

Data owners usually contract with vendors to build data systems or subscribe to cloud-based services to support mission capabilities. Outsourcing this work is a good government practice since it leverages the expertise of the private sector to support the needs of the public. However, major issues arise when the public sector organization doesn’t have access to its own data when it’s housed in a proprietary vendor-housed or cloud-based data system. Oftentimes, vendors require the public sector organization to use its reporting infrastructure to access data.

This creates significant inefficiencies when the organization wants to leverage the data asset for integration with other datasets. Virginia experienced this during the COVID-19 pandemic response when some data assets were not available programmatically, but rather, required a human to log on to the application, run a custom report, and download data on a daily basis to support critical business functions. Ideally, there should be direct access to data either through direct database connections or API calls, but unfortunately, we had to implement Robotic Process Automation to streamline the data extraction process. We also had some data sources that automatically distributed data reports via email, but found that the data structure of the extract changed frequently without advance notice creating issues in the data ingestion process.

Unfortunately, this is not a technical issue, but an issue in the language used in technology procurement contracts and their enforcement. In Virginia, the procurement contract language includes references to “works”, “work products”, “intellectual property”, and “data”, but is focused primarily on granting use of “Commonwealth Works” in a limited, non-exclusive, non-transferable, and royalty-free license to the vendor while also limiting its use to only those processes directly associated with the contracted services.

Moving forward, technology procurement contract language should include provisions for the public sector entity to have direct access to all data stored and managed by the vendor either through APIs or some other programmatic means. Our ability to leverage public sector data for the common good should not be limited by having human resources dedicated to manually moving data from one platform to another. Many, if not all, data transfer processes should be fully automated and programmatically controlled to minimize human involvement and the potential for human error.

The articles from these contributors are based on their personal expertise and viewpoints, and do not necessarily reflect the opinions of their employers or affiliated organizations.
Top