Hybrid approach combines topic and graph embeddings for legal document clustering.
problem Challenges in classifying legal texts due to domain-specific language and limited labeled data.
method Combines unsupervised topic and graph embeddings with a supervised model.
result Improves clustering quality over text-only or graph-only embeddings.
This paper improves NEL for legal documents using transfer learning.
problem Linking named entities in legal documents.
method Transfer learning approach applied to legal datasets.
result Improved F1-score of 98.90% and 98.01% on legal datasets.
Deep learning improves legal document translation, summarization, and classification.
problem Data scarcity in legal document processing.
method Multi-task deep learning to leverage transfer learning.
result Multi-task DL outperformed state-of-the-art results in all tasks.
Bi-LSTM classifies legal documents for Brazil's supreme court.
problem Clogging of Brazil's supreme court due to high volume of lawsuit cases.
method Used a Bidirectional Long Short-Term Memory (Bi-LSTM) network.
result Successfully classified legal documents for efficient case allocation.
Paper proposes GSSNMF for legal document classification and topic modeling.
problem Lack of methods that can both classify and model topics with guidance.
method Guided Semi-Supervised Non-negative Matrix Factorization (GSSNMF).
result Improves both classification accuracy and topic coherence.
Machine Learning community is recently exploring the implications of bias and fairness with respect to the AI applications. The definition of fairness for such applications varies based on their domain of application. The policies governing the use of such machine learning system in a given context are defined by the c…
LexNLP is an open source Python package focused on natural language processing and machine learning for legal and regulatory text. The package includes functionality to (i) segment documents, (ii) identify key text such as titles and section headings, (iii) extract over eighteen types of structured information like dis…
Case Law has a significant impact on the proceedings of legal cases. Therefore, the information that can be obtained from previous court cases is valuable to lawyers and other legal officials when performing their duties. This paper describes a methodology of applying discourse relations between sentences when processi…
Paper presents LLM-enhanced contract metadata extraction.
problem Automatic detection and annotation of legal clauses in contracts.
method Integration of publicly available and proprietary datasets with advanced LLM methodologies.
result Substantial improvements in clause identification accuracy and efficiency.
Computable contracts simplify financial transactions and reduce legal costs.
problem Difficulty in querying, executing, and analyzing text-based financial contracts.
method Develop a Contract Definition Language and illustrate use cases.
result Substantial improvements in customer experience and cost reduction.
Classifies privacy policy segments for better user understanding.
problem Difficulty in understanding privacy policies due to legal jargon.
method Uses machine learning and deep learning techniques to classify privacy policy segments.
result Identifies data practices in privacy policies for better user comprehension.
This paper develops a structural credit risk model to characterize the difference between the economic and recorded default times for a firm. Recorded default occurs when default is recorded in the legal system. The economic default time is the last time when the firm is able to pay off its debt prior to the legal defa…
This paper analyzes crypto white papers under MiCAR, highlighting NLP's role.
problem Regulatory changes in crypto white papers under MiCAR.
method Survey of existing NLP applications, analysis of MiCAR changes.
result NLP can assist in regulatory compliance and white paper analysis.
The paper highlights legal misunderstandings in ML fairness definitions.
problem Misalignment between ML fairness definitions and legal concepts.
method Examples and comparative analysis of legal and ML terminology.
result Both communities need to learn from these tensions.
Paper formalizes anti-discrimination law in automated systems.
problem Algorithmic discrimination in legal contexts.
method Decision-theoretic framework grounded in UK anti-discrimination law.
result Introduced 'conditional estimation parity' metric for ML fairness.
Bayesian CNN estimates uncertainty in bone age prediction.
problem Uncertainty quantification in age estimation models.
method Variational Inference for Bayesian CNNs.
result Model uncertainty distinguished from data uncertainty.
Causal discovery algorithms can help generate legal arguments.
problem Leveraging causal discovery algorithms in legal decision-making.
method Prepared a legal dataset, annotated with 17 legal concepts, applied causal discovery algorithms, and quantified degrees of belief.
result Some causal relationships help generate viable legal arguments.
This paper predicts legal proceedings status using NLP and machine learning.
problem Classify Brazilian legal proceedings into archived, active, and suspended categories.
method Combined NLP techniques with machine learning to classify legal proceedings sequences.
result Achieved maximum accuracy of 93% and top average F1 Scores of 89% (macro) and 93% (weighted).
We analyze the practical consequences of the bilateral counterparty risk adjustment. We point out that past literature assumes that, at the moment of the first default, a risk-free closeout amount will be used. We argue that the legal (ISDA) documentation suggests in many points that a substitution closeout should be u…
Study identifies and measures biases in legal case data.
problem Addressing representation biases and sentencing disparities in legal case data.
method Utilizes two regression models: a baseline and a fair judge model.
result Quantifies biases across demographic groups in criminal data from Cook County (Illinois).
Research creates a taxonomy to bridge AI security and regulatory gaps.
problem Disciplinary disconnect between technical and legal teams in AI risk assessment.
method Developed an AI System Threat Vector Taxonomy with 9 domains and 53 sub-threats.
result Empirically validated and aligned with ISO/IEC 42001 controls and NIST AI RMF functions.
Formulates LGFO to measure fair ML systems using legal signals.
problem Formally incompatible measures of unfairness in ML systems.
method Uses legal signals to measure social cost of unfairness.
result LGFO aligns with societal view of unfairness.
Insiders camouflage trading to balance wealth and stealth, avoiding legal penalties.
problem Legal penalties and insider trading among liquidity traders.
method Kyle-type model with a diverse spectrum of prosecution schemes.
result Existence and uniqueness of equilibria for large populations, with a stealth index revealing trading scale.
Smart Close-out Netting aims to automate close-out netting processes.
problem Inefficiencies in close-out netting processes for financial institutions.
method Standardisation and automation of legal and regulatory processes using a data-driven framework and controlled natural language.
result Standardisation and automation can improve close-out netting processes for prudentially regulated financial institutions.
Insider trading is reduced when penalized, affecting expected penalties in a non-monotone way.
problem Reducing insider trading behavior when insiders face legal penalties.
method Characterized via a backward stochastic differential equation (BSDE) with a non-linear operator.
result The insider's expected penalties are non-monotone in the fee structure and determined by relative entropy.
This research uses deep learning to automatically classify UN resolutions.
problem Manual labeling of UN documents is too time-consuming.
method Utilizes pre-trained deep learning models without traditional training.
result Shows effectiveness in classifying UN resolutions by SDGs.
Deep learning model tackles illegal comments with promising results.
problem Lack of labeled data in legal domain.
method Multi-task deep learning model for classification.
result Promising results in classifying illegal comments.
Federated learning improves bioinformatics by sharing data legally.
problem Lack of access to diverse data in bioinformatics.
method Combines data from multiple institutions legally.
result Federated learning accelerates clinical discovery and robust exploration.
What does it mean for an algorithm to be biased? In U.S. law, unintentional bias is encoded via disparate impact, which occurs when a selection process has widely different outcomes for different groups, even as it appears to be neutral. This legal determination hinges on a definition of a protected class (ethnicity, g…
ClauseLens uses reinforcement learning to price reinsurance treaties transparently and auditably.
problem Opaque and difficult-to-audit reinsurance treaty pricing practices.
method ClauseLens models treaty pricing as a Risk-Aware Constrained Markov Decision Process (RA-CMDP), incorporating legal clauses and generating interpretable explanations.
result ClauseLens reduces solvency violations and improves tail-risk performance, achieving 88.2% accuracy in clause-grounded explanations.
Over the last 23 years, the U.S. Securities and Exchange Commission has required over 34,000 companies to file over 165,000 annual reports. These reports, the so-called "Form 10-Ks," contain a characterization of a company's financial performance and its risks, including the regulatory environment in which a company op…
New approach to counterfactual reasoning avoids demographic interventions.
problem Limitations of traditional counterfactual reasoning in AI systems.
method Backtracking counterfactual approach instead of interventional.
result Allows addressing social concerns without demographic interventions.
GPT learns a causal world model from token predictions, validated in game sequences.
problem Does GPT implicitly learn a causal world model from token predictions?
method Derived a causal interpretation of GPT's attention mechanism and proposed zero-shot causal structure learning.
result GPT can generate legal next moves with high confidence for sequences with encoded causal structures, but fails for illegal moves.
Framework uses deep learning to analyze large documents and identify their logical structure.
problem Analyzing large, multi-themed documents with diverse topics.
method Deep learning techniques to model and extract logical and semantic structure.
result Framework effectively identifies and classifies different sections of documents.
Law responds to adversarial machine learning threats.
problem Adversarial machine learning attacks and their legal implications.
method Scenarios and legal analysis of adversarial ML attacks.
result Some attacks are more likely to result in liability.
Offline RL algorithms protect privacy while learning from sensitive data.
problem Learning sensitive data in financial, legal, and healthcare applications.
method Differential privacy guarantees for offline RL algorithms.
result Provable prevention of privacy risks with strong learning bounds.
Examines AI regulation in finance, highlighting risks and gaps in current laws.
problem Rapid AI adoption in finance introduces risks and compliance challenges.
method Reviews current legislation, industry guidelines, and real-world use cases.
result Need for adaptive, technology-neutral policies to balance innovation and consumer protection.
A new method for document network embedding interprets and generalizes well.
problem Lack of interpretability and generalization to new documents in existing methods.
method Introduces Topic-Word Attention (TWA) and Inductive Document Network Embedding (IDNE) to generate document representations.
result Achieves state-of-the-art performance on various networks and produces meaningful representations.
A new model CDTM improves text classification by concentrating document topics.
problem Unsupervised text classification with diverse topic distributions.
method Imposes an exponential entropy penalty on document topic distribution to encourage concentration.
result More coherent topics and concentrated, sparse document-topic distributions.
New document embedding method finds more relevant tech docs.
problem Computing document similarities for technical content.
method Hybrid approach using graph techniques to extract keyphrases and score sentences.
result Proposed methods outperform baselines up to 27% in NDCG.
Tokenized RWAs face liquidity issues despite promising markets.
problem Low trading volumes and limited investor participation in tokenized assets.
method Empirical analysis of tokenized real estate, private credit, and treasury funds.
result Most tokenized assets exhibit low transfer activity and limited secondary trading.
Improved self-supervised learning for document images.
problem Performance of self-supervised pre-training on document images is poor.
method Proposed context-aware alternatives and a novel multi-modal method.
result Novel method outperforms other self-supervised methods on document image classification.
MarlRank uses multi-agent reinforcement learning to improve document ranking.
problem Neglecting mutual information among documents in ranking models.
method Formulated as a multi-agent Markov Decision Process (MDP), each document predicts relevance considering its own and similar documents features and actions.
result Significant performance gains over state-of-the-art baselines on LETOR benchmark datasets.
DocParser parses document structures from renderings like PDFs and scans.
problem Parsing complete hierarchical document structures from renderings.
method End-to-end system with novel weak supervision approach.
result Significant improvement in document structure parsing performance.
Browsing and finding relevant information for Bangladeshi laws is a challenge faced by all law students and researchers in Bangladesh, and by citizens who want to learn about any legal procedure. Some law archives in Bangladesh are digitized, but lack proper tools to organize the data meaningfully. We present a text vi…
CCAligned creates a massive web document dataset for cross-lingual research.
problem Identifying comparable documents across different languages.
method Using URL signals to label web documents and mining Common Crawl corpus.
result Release of a dataset with over 392 million URL pairs from 8144 language pairs.
Smart contracts create digital financial derivatives.
problem Creating a new digital financial derivative contract.
method Applied existing smart contract technologies to develop two prototypes.
result Demonstrated feasibility of digital financial derivatives on centralized and DLT platforms.
Understanding large, structured documents like scholarly articles, requests for proposals or business reports is a complex and difficult task. It involves discovering a document's overall purpose and subject(s), understanding the function and meaning of its sections and subsections, and extracting low level entities an…