Análisis de sesgos en un modelo de Procesamiento de Lenguaje Natural para clasificar proyectos Capital Semilla de Sercotec.
Loading...
Date
2024
Authors
Journal Title
Journal ISSN
Volume Title
Publisher
Universidad de Concepción
Abstract
El Servicio de Cooperación Técnica (Sercotec) presenta un importante desafío: poder evaluar todas las postulaciones al Fondo Capital Semilla, el cual entrega recursos a nivel nacional para emprendimientos en etapa temprana. Para abordar esta tarea no solo es necesario crear un mecanismo que permita eficientizar el proceso de revisión de los proyectos, sino también asegurarse de que este sea imparcial, no reproduciendo sesgos o discriminación entre los postulantes.
Para dar solución a esta problemática, se entrenan dos modelos de aprendizaje automático utilizando el modelo de Procesamiento de Lenguaje Natural BETO, llevando a cabo uno para el ítem de Clientes y otro para el ítem de Oferta de Valor del modelo de negocios Canvas. El objetivo es evaluar de manera automatizada las respuestas de los postulantes y calificarlas según una escala de puntaje diseñada por Sercotec. Además, se analizan dos tipos de sesgos que podrían estar presentes en los modelos de aprendizaje automático.
Por un lado, se analiza el sesgo social, cuando los modelos generan resultados favorables o desfavorables para ciertos grupos específicos. Se generan cambios de género en las
respuestas de uno de los ítems para analizar las discrepancias en la clasificación del modelo al realizar este proceso. Por otro lado, se estudia el sesgo de etiqueta, cuando los datos podrían estar incorrectamente etiquetados, reflejando alguna tendencia o imparcialidad, entregando resultados con los mismos sesgos. Se estudian las tendencias de los evaluadores al calificar proyectos, detectando aquellos que asignan una mayor o menor proporción de notas en comparación con sus compañeros en algún ítem de la postulación.
Para el primer tipo de sesgo se obtiene una diferencia en el 5% de las respuestas, lo cual, aunque es un porcentaje bajo, debería ser idealmente nula. Para el segundo sesgo, se
determina que estas tendencias, sumadas a la cantidad de proyectos que se evalúan, afectan considerablemente el resultado de la evaluación.
The Technical Cooperation Service (Sercotec, by its initials in Spanish) presents an important challenge: to evaluate all applications for the Seed Capital Fund, which provides resources nationwide for early-stage ventures. To address this task, it is necessary not only to create a mechanism to streamline the project review process but also to ensure that it is impartial and does not reproduce bias or discrimination among applicants. To solve this problem, two machine learning models are trained using the BETO Natural Language Processing model, implementing one for the Customers item and another for the Value Proposition item of the Canvas business model. The objective is to evaluate in an automated way the answers of the applicants and score them according to a scoring scale designed by Sercotec. Additionally, two types of biases that could be present in machine learning models are analyzed. Firstly, social bias is analyzed, where the models generate favorable or unfavorable results for specific groups. Gender changes in the responses of one of the items are generated to analyze discrepancies in the model's classification when performing this process. Secondly, label bias is studied, where data could be incorrectly labeled, reflecting some tendency or partiality, resulting in outputs with the same biases. The tendencies of the evaluators when grading projects are studied, detecting those who assign a higher or lower proportion of marks compared to their peers in any item of the application. For the first type of bias, a difference is observed in 5% of the responses, which, although a low percentage, should ideally be zero. For the second bias, it is determined that these tendencies, combined with the number of projects evaluated, significantly affect the evaluation results.
The Technical Cooperation Service (Sercotec, by its initials in Spanish) presents an important challenge: to evaluate all applications for the Seed Capital Fund, which provides resources nationwide for early-stage ventures. To address this task, it is necessary not only to create a mechanism to streamline the project review process but also to ensure that it is impartial and does not reproduce bias or discrimination among applicants. To solve this problem, two machine learning models are trained using the BETO Natural Language Processing model, implementing one for the Customers item and another for the Value Proposition item of the Canvas business model. The objective is to evaluate in an automated way the answers of the applicants and score them according to a scoring scale designed by Sercotec. Additionally, two types of biases that could be present in machine learning models are analyzed. Firstly, social bias is analyzed, where the models generate favorable or unfavorable results for specific groups. Gender changes in the responses of one of the items are generated to analyze discrepancies in the model's classification when performing this process. Secondly, label bias is studied, where data could be incorrectly labeled, reflecting some tendency or partiality, resulting in outputs with the same biases. The tendencies of the evaluators when grading projects are studied, detecting those who assign a higher or lower proportion of marks compared to their peers in any item of the application. For the first type of bias, a difference is observed in 5% of the responses, which, although a low percentage, should ideally be zero. For the second bias, it is determined that these tendencies, combined with the number of projects evaluated, significantly affect the evaluation results.
Description
Tesis presentada para optar al título de Ingeniero/a Civil Industrial.
Keywords
Procesamiento de lenguaje natural (Lingüística, Sesgo estadístico, Procesamiento de datos