Addressing image-classification model quality degradation when access to new data is unavailable under data drift or concept drift.
The problem
The vendor’s Base model becomes outdated due to data drift or concept drift, but the Client’s data is unavailable to the Vendor. Legal or organizational restrictions — HIPAA, GDPR, industry regulators — prohibit transferring sensitive data from the Client to the Vendor. The classic centralized retraining cycle is either legally blocked or substantially impeded.
Microsoft reports that machine learning models can lose over 40% of their accuracy within a year if data drift is not accounted for. Meanwhile, obtaining new representative data and running multiple retraining iterations takes from several weeks to months per cycle.
The solution and outcome
FedTuna implements federated fine-tuning between the Vendor and the Client. Only weights, gradients, and quality metrics are exchanged — never the Client’s source data. The Client labeled 100 images (50 per class). After 5,000 fine-tuning iterations, the model significantly improved its quality on the Client’s data while preserving its quality on the Vendor’s validation dataset without noticeable degradation.
The technical component of drift response was reduced from 24 weeks to 6 days. The operating cost of refreshing the model was reduced by 50%. The contractual model did not change.
Who is involved and what they bring.
The Vendor
A supplier of an ML-based IT solution for image classification, for whom access to the Client’s data is necessary to maintain high ML model quality, but is restricted or impossible due to legal, organizational, or other reasons.
The Client
A user of the Vendor’s ML solution: a healthcare organization, a bank, an insurance company, a fintech company, etc. In the scope of the task, this is the owner of a protected on-premise infrastructure that operates the Vendor’s ML solution in a production environment and accumulates new operational data that is unavailable to the Vendor.
The Product
A solution from FedTuna — a federated fine-tuning platform that is embedded into the process of training and refreshing the Vendor’s models. The customer and principal consumer of the Product is the Vendor.
The Vendor already has a Base model: a ready-to-use image classification model that the Vendor delivers to its Clients either separately or as part of its own software. The Vendor delivers its software on-premise, i.e., once the software has been handed over to the Client, the Vendor no longer has direct access to it or to the Client’s data.
At the same time, technical communication channels may be used between the infrastructures of the Client and the Vendor — for example, for updating the solution and for transferring privacy-preserving data: transaction metrics, model accuracy figures, and dataset statistics. However, such channels must not transmit information containing confidential data, including personal data and information that constitutes a commercial secret.
The importance of the Product stems from the requirements for protecting sensitive information that are typical of a B2B scenario. These include:
- Regulatory constraints;
- Requirements set by the Client's information security team;
- Internal data access policies;
- Obtaining the data owner's consent for the use of the data outside the Client's perimeter;
- Contractual obligations toward end users;
- Requirements for validating and certifying medical software (FDA 510(k)/De Novo, CE marking under MDR, IEC 62304), under which any ML model update affecting a clinical decision may be classified as a change to a medical device and may require repeated regulatory assessment.
The Base model becomes outdated. The Client's data is unavailable.
Core premise: the Vendor has a Base model.
Key problems
- The Base model becomes outdatedMicrosoft reports that machine learning models can lose over 40% of their accuracy within a year if data drift is not accounted for.microsoft.com/en-us/research/…/MLSYS2022.pdf (loses quality) due to data drift or concept drift.Data drift is a change in the input data. Concept drift is a change in input-output relationships. Both often happen simultaneously.evidentlyai.com/ml-in-production/data-drift
- The Client’s data is unavailable to the Vendor. Moreover, the Client either does not label the data at all or labels it only partially.
- Creating or finding data with a profile similar to the Client’s data is a long and costly process with no guarantee of success.
- Synthetic data often delivers limited benefit and does not always allow the real characteristics of the Client’s data to be reproduced
- To assess the result, the centralized model must be updated and its behavior re-evaluated in a closed environment. In some cases, the number of such iterations may reach dozens and stretch over months.
- In medical imaging, assessing model behavior requires not only technical validation on metrics (F1, AUC), but also clinical validation involving domain specialists (radiologists, pathologists) and, in some cases, a reader study, which further increases the duration and cost of the cycle.
Causes of data drift and concept drift
- Emergence and market spread of new devices, cameras, sensors, and other data sources.
- Changes in software used to capture, transmit, and process images, video, text, and other data.
- Changes in the composition and structure of the observed population, including shifts in demographic, age-related, physiological, and other characteristics. In medical imaging, this may manifest as changes in the patient population by comorbidities, distribution of disease stages, ethnic composition, body habitus, and other clinically significant parameters.
- Changes in user scenarios, in the conditions under which data is collected, and in the context in which the system is used: data may be collected indoors or outdoors, at various times of day, under insufficient, excessive, or unstable lighting, and in other changing conditions.
- Changes in business processes, in the rules for interpreting results, and in the criteria for assigning objects to target classes.
- Emergence of new types of objects, events, anomalies, and, in the case of cybersecurity threats, that were previously absent from the training data or were only weakly represented in it.
- Changes in the behavior of users, external participants, or malicious actors, including adaptation to the system and the appearance of new ways of circumventing it.
The classic path and its limitations
- Obtain new, up-to-date data in one of the following ways:
- •Receive up-to-date data from the Client.
- •Find open data on the internet.
- •Purchase data legally.
- •Create synthetic data.
- •Generate the data in-house.
- Label the obtained data.
- Fine-tune or retrain from scratch the Vendor’s model on the centralized data.
- Deliver the updated version of the Vendor’s model to the Client.
This approach is widely used in practice, but it has a number of limitations. In the scenario under consideration, it is either inapplicable or substantially impeded.
Where this may be relevant
The problems described are characteristic, in particular, of healthcare, where new subtypes may emerge within an already known class over time, or the composition of visual manifestations may change. For example:
- Mammography: within the class of malignant lesions, the model may increasingly encounter cases with different morphological subtypes or a different visual presentation.
- Diabetic retinopathy: within the positive class, the composition of visual manifestations may change, including microaneurysms, hemorrhages, and other diagnostically significant features.
- Histopathology: among malignant cases, the model may begin to encounter more frequently rare histological variants or variants that were not previously represented in the training set.
- Dermatoscopy: in tasks where the positive class combines malignant skin lesions, the ratio of different subtypes of melanoma and other malignant neoplasms, as well as the distribution of cases across skin phototypes and anatomical zones, may change over time. As a result, such a class may include visually distinct subtypes that were absent or weakly represented in the original training sample.
Similar situations also arise in other industries when new subtypes, patterns, or anomalies that are not represented in the training data appear within an already known class — for instance, in cybersecurity and fraud prevention. In all of these scenarios, the common pattern is the same: the Client gets new data or new object subtypes that are unavailable to the Vendor but that are critical to maintaining the model’s quality.
How the FedTuna Product works in practice.
The FedTuna Product implements federated fine-tuning between two Parties: the Vendor and its Client.
Main scenario
- The Product is embedded into the Vendor’s software that is delivered to the Client. FedTuna itself does not become a party to the interaction. The interaction takes place between the Vendor and the Client using the Product.
- The Parties (the Vendor and the Client) establish a connection (with TLS encryption) between their servers over the gRPC protocol.
- The Parties prepare the artifacts required for joint fine-tuning: a project, a model, a model version, and the training and validation data. The bulk of the preparation is carried out on the Vendor’s side, while the actions required from the Client remain minimal.
- The Parties launch the replication process, after which they synchronize the entities of the Product at a defined frequency. The Client’s involvement is reduced to a minimal set of actions (uploading data for fine-tuning, providing the Product with access to it), whereas the bulk of process control is carried out on the Vendor’s side. Data preparation is performed locally inside the perimeter of each Party, while the fine-tuning itself is performed jointly.
During model fine-tuning
- Only weights, gradients, and quality metrics are exchanged between the Parties, never the Client’s source data. In a deep neural network, gradients are formed as the result of a long chain of computations and nonlinear transformations. State-of-the-art methods for reconstructing source data from gradients yield only partially approximate reconstructions even in laboratory conditions.
- Control is enforced so that the metrics of the fine-tuned model are no worse than the metrics of the previous version of the model. The metrics in question are those computed on the Vendor’s validation dataset.
- At each fine-tuning step, computations are performed both on the Vendor’s side and on the Client’s side, after which the results of the computations are aggregated. This scheme makes it possible to adapt the model to new Client data without losing quality on the Vendor’s original data.
Labeling conditions on the Client’s side
- Initially, labeling may be absent because the Client does not have the necessary resources or expertise, while labeling by the Vendor is not possible due to the prohibition on transferring sensitive data between the parties.
- Various scenarios for the appearance of labeled data are possible: labeling of historical data on the Client’s side, manual labeling of a limited sample by the Client, and gradual accumulation of labeled cases through the inference of the existing model.
- In a practical scenario, the Client may, for example, label 50–100 cases, which can be enough to launch a limited fine-tuning cycle or to perform an initial check of hypotheses. It should be borne in mind that in medical imaging, labeling requires qualified specialists (radiologists, pathologists), and to confirm the clinical significance of the improvements, the validation sample, as a rule, must be substantially larger.
- Labeled data may appear naturally, for example, via end-user requests to the support team or when there are clearly expressed signs of the target class, even in the absence of an organized, systematic labeling process.
- In medical imaging, “natural labeling” via support requests or clinician feedback is one of the possible sources, but it is subject to systematic bias: predominantly obvious errors, not borderline cases, come into focus. To obtain representative labeling, structured procedures involving domain specialists are required.
Distribution of data and artifacts
The Vendor has: the model; datasets used for training and validating the model and for assessing its quality metrics.
The Client has: the Vendor’s Base model, supplied as part of the Vendor’s IT solution; a set of new, up-to-date data generated in the course of operating the Client’s IT systems; historical or current data labeled in some way.
Computational resources (experiment configuration)
| Component | Specification |
|---|---|
| CPU | 16 vCPU |
| RAM | 32 GB |
| SSD | At least 50 GB of free space |
| GPU | NVIDIA Tesla T4, 16 GB VRAM |
In the experiments, the same configuration was used on each Party. The Product does not require a GPU on the Client’s side in order to operate, but, all other things being equal, in a configuration without a GPU on the Client’s side, training time may roughly double compared with the same configuration with a GPU.
Interaction workflow
ML SOLUTION VENDOR
Runs projects · owns models
VENDOR'S CLIENT
Syncs objects · provides data
Perform configuration
modify configuration files
Launch docker containers
standalone or as part of vendor's product
Launch the product
import and initialize the Python package
Export the vendor's public key
and send it to the client
Establish a connection
with the client
import the client's binary file
pubkey
connect
Export the client's public key
and send it to the vendor
Establish a connection
with the vendor
import the vendor's binary file
Launch the automatic synchronization process
Periodically sync the list of objects with the vendor
Create a project, a model, and a model version
Wait for the project list to be synchronized
Prepare and upload training data
into the project
Launch model version
fine-tuning
Participate in model fine-tuning
Stop fine-tuning and obtain its results
weights
gradients
model metrics
Wait for the fine-tuning session list to be synchronized
Participate in model fine-tuning
Figure 1. Interaction workflow between the Vendor and the Client.
What was simulated, and how.
Within the pilot project, the following problem was reproduced: data drift or concept drift caused by the appearance of a new class of objects, which led to a rise in the number of false positives. A binary image-classification model was used for the experiment.
Description of the experiment
Simulated situation: a growing volume of new-class data on the Client’s side, caused by the appearance of images of a new type or a new class of objects that were absent from the original training sample.
Vendor’s Base model:
- Was trained only on the Vendor’s data, which did not contain images of the new type;
- Was used as a starting point for fine-tuning on new Client data that was unavailable to the Vendor.
Experiment goal: to assess the impact of federated fine-tuning on the model’s ability to simultaneously:
- Adapt to images of the new type;
- Preserve quality on the Vendor’s data.
This case reflects data drift. In the Client’s production environment, a new class or subtype of anomalies appears that was not represented in the training of the Vendor’s Base model. Such a scenario is characteristic of any binary image-classification task in which the structure of the negative class changes over time faster than a centralized model update is feasible: healthcare, fintech, industrial quality control, information security, and others.
Dataset characteristics
Anomalies were singled out from the overall set of images. In this case, anomalies refer to images of the new type that were completely excluded from the training sample of the Base model and were therefore classified incorrectly by that model. The Client’s training and test datasets were assembled on the basis of such images. The remaining images were used to train the model on the Vendor’s side. Then the model was fine-tuned using the Client’s data.
| Parameter | Vendor side | Client side |
|---|---|---|
| Training images | >250,000 | 100 (50 per class) |
| Validation / test images | >18,000 | >3,000 (test set) |
| Batch size | 32 | 8 |
| Learning rate | 5e-5 | 5e-5 |
| Gradient weight at aggregation | 0.8 | 0.2 |
| Anomaly selection (client train) | Highest model uncertainty: score ∈ [0.1, 0.3] | |
The embedding space tells the story.
Training speed results are deliberately not published, because they depend on a large number of factors: the network architecture of the Parties, network latency and channel throughput, the information-security tools used, the available computational resources, as well as the model size and the volume of data for fine-tuning.
Embedding space of the Base model
The original embeddings have a dimensionality of 1536, so their distribution cannot be visualized directly in a two-dimensional space. Below the dimensionality of the embeddings was reduced to two components using the t-SNE method.
Both classes and Anomalies
Class 1 and Anomalies
Embedding space of the Base model fine-tuned with the Product
Both classes and Anomalies
Class 1 and Anomalies
Data drift
As part of the experiment, a check for data drift was performed. The methodology of the check follows. An analysis of data drift based on the Euclidean distance between the averaged embeddings of the reference and the current dataset, when anomalies’ embeddings are mixed into the current dataset (in 2% steps), produced the following results:
Visually, data drift is barely detectable. However, if a data-drift analysis is performed by mixing anomalies only into the embeddings of class-1 images, a clearly pronounced tendency for the distance to increase appears:
This indirectly confirms the presence of data drift on class 1. To obtain additional confirmation or refutation of the hypothesis of data drift, an analysis was performed using Evidently, which showed the following:
When analyzing data drift using Evidently:
- The dimensionality of the embeddings was first reduced to 3 components.
- The following parameters were appliedParameters for Evidently data drift detection.docs.evidentlyai.com/metrics/customize_data_drift:
- •The share of drifting columns as a condition for Dataset Drift: 30%.
- •Data drift detection method: Wasserstein distance (normed) with threshold 0.07.
Dataset Drift
Dataset Drift is detected. Dataset drift detection threshold is 0.3
3
Columns
1
Drifted Columns
0.333
Share of Drifted Columns
Drift is detected for 33.333% of columns (1 out of 3).
emb_2
Detected
Type:num
Drift Score:0.077561
Stat Test:Wasserstein distance (normed)
Reference Distribution
Current Distribution
emb_0
Not Detected
Type:num
Drift Score:0.051847
Stat Test:Wasserstein distance (normed)
Reference Distribution
Current Distribution
emb_1
Not Detected
Type:num
Drift Score:0.02867
Stat Test:Wasserstein distance (normed)
Reference Distribution
Current Distribution
Dataset Drift detected. Threshold: 30% of columns drifting. 1 out of 3 columns (33.3%) drifted.
The point of the analysis performed boils down to checking whether the new data containing anomalies starts to differ statistically from the distribution on which the Base model was trained. Taken together, the results listed above make it possible to assert with a high degree of confidence that data drift is present.
Model metrics
The number of fine-tuning iterations was 5,000. On the charts, the horizontal line marks the level of the metric corresponding to its value before fine-tuning of the Base model.
Interpretation: On the Client’s data, the model’s quality improved noticeably. On the Vendor’s validation dataset, no noticeable degradation of the model is observed. The fine-tuned model has adapted to the new data while retaining its ability to perform on the Vendor’s original data with no noticeable degradation.
Federated fine-tuning as a practical mechanism for model refresh.
The experiment conducted shows that federated fine-tuning can be a practical way to refresh an ML model in conditions where new data appears only on the Client’s side and cannot be transferred to the Vendor.
In the case under consideration, the Vendor’s Base model was not trained on the new type of anomalies that appeared in the Client’s production environment. After federated fine-tuning the model significantly improved its quality on the Client’s data while preserving its quality on the Vendor’s validation dataset without noticeable degradation.
This means that the approach makes it possible to solve two tasks simultaneously:
- Adapt the model to new operating conditions;
- Keep its original performance on the Vendor’s side under control.
The practical implication of the result is that even a limited amount of data on the Client’s side may be sufficient to start a new cycle of model adaptation. In the demonstrated scenario a small number of labeled incidents was used on the Client’s side, while the bulk of the source data and the model’s quality control remained on the Vendor’s side. This regime is especially important for B2B on-premise scenarios in which transmitting the Client’s images, video or other sensitive data is not possible for legal, organizational or other reasons. In these conditions federated fine-tuning serves not as a research alternative, but as a practically realizable mechanism for refreshing the model so that it can subsequently be operated under real-world conditions.
Applicability across industries
The applicability of the approach is not limited to one particular domain. It may be in demand in any image- or video-classification task in which model quality deteriorates due to the appearance, on the Client’s side, of new data that was not previously represented. In healthcare, this may be linked to new subtypes of pathology or changes in the visual composition of the positive class; in antifraud, to new types of attacks; in industrial quality control, to new types of defects; in information security and content moderation, to new patterns of violations or anomalies. The common feature of all such scenarios is the same: new relevant data appears at the Client’s side before the Vendor has an opportunity to update the model centrally.
Medical imaging
New subtypes may emerge within an already known class — mammography (different morphological subtypes), diabetic retinopathy (changes in microaneurysms, hemorrhages), histopathology (rare histological variants), dermatoscopy (different subtypes of melanoma across skin phototypes and anatomical zones).
New types of attacks
In cybersecurity and fraud prevention, new types of objects, events, anomalies that were previously absent from the training data or were only weakly represented in it appear continuously as adversarial actors adapt to the system.
New types of defects
New defect types emerge within an already known class when equipment, materials, or production conditions change. The client accumulates new operational data unavailable to the vendor but critical to maintaining model quality.
New patterns of violations
New patterns of violations or anomalies appear as the context in which the system is used changes. In all such scenarios the common pattern is the same: the Client gets new data or new object subtypes that are unavailable to the Vendor but critical to maintaining model quality.
Note on regulated industries (medical imaging, SaMD): In regulated industries, primarily in medical imaging, the practical adoption of federated fine-tuning requires that additional constraints be taken into account: a model used to support clinical decisions may fall under regulation as SaMD (Software as a Medical Device), and any update to it must be assessed within the applicable regulatory framework (FDA, MDR/IVDR). This does not negate the value of the approach, but it does affect the permissible frequency of updates and requires that a process for regulatory support be put in place. The figures cited for timelines and savings are characteristic of non-regulated industries.
A separate conclusion from the experiment’s results is that federated fine-tuning can be used not only as a response to an isolated incident but also as a controlled process for planned model refresh. As new data accumulates on the Client’s side, the approach makes it possible to gradually adapt the model to the changing operating environment without removing data from the Client’s infrastructure and without fully restarting the classic centralized retraining cycle. This makes the product especially valuable where the classic path is too slow, too costly, or organizationally too difficult.
The classic model-refresh cycle takes from several weeks to months and costs around $20,000 per client per year.
The classic model-refresh cycle: receive data from the Client, find or create equivalents, label them, retrain centrally, deliver the update. It takes from several weeks to several months and requires the involvement of at least two parties at each stage. At the same time, the iterative cycle of “update, evaluate in a closed environment, repeat” may be reproduced dozens of times.
The total cost of a single refresh cycle includes the time that ML engineers spend on retraining and validation, as well as the legal and organizational work involved in transferring the data. On top of that, there are direct losses from the degrading model in the period between updates: rising rates of false rejections or growing numbers of missed cases translate directly into operating costs and reputational damage. In medical imaging, the consequences may be more severe: a missed pathology (a false negative) carries a direct risk to patient health, while a false positive risks unnecessary invasive procedures.
Federated fine-tuning eliminates the most costly components of this cycle: the data does not leave the Client’s perimeter, labeling is required only for a small fraction of the incoming flow, and model adaptation happens iteratively and continuously, without stopping the production environment.
| Metric | Classic cycle | Federated fine-tuning |
|---|---|---|
| Single refresh cycle cost | ~$10,000 | ~50% reduction |
| Annual cost per client (2× / year) | ~$20,000 | ~$10,000 |
| Reaction time to drift | 24 weeks | 6 days |
| False rejection rate increase | 5–30 pp between updates | Continuous adaptation |
| Raw data transferred | Yes | Zero bytes |
| Change to contractual model | Required | None |
Cost includes engineering time, labeling and compute
In medical imaging, the final time-to-deployment is determined not only by the speed of fine-tuning but also by the duration of the regulatory assessment of the updated model. Nevertheless, the federated approach reduces precisely the technical component of the cycle, which under the classic scenario also accounts for a significant share of the total time.