What ARGIRA is

ARGIRA is a research and accessibility initiative that explores new ways of representing information through technology, including sonification approaches designed to create richer experiences for people with visual impairments.

Ongoing research · 2026

ARGIRA
Sonification
and Perceptual
Metrics

Ranero García, Jose
Zenodo · CC BY-NC 4.0

New · ARGIRA v1.4.7 · Alt-text uncertainty · Aug 17, 2026

Can the visual structure of an image predict how it perceptually sounds? ARGIRA is a research project that answers this question through systematic negative results — mapping the limits of what standard visual representations can and cannot explain about sound perception.

This page is the index and summary of the ARGIRA study series (I–XII), published on Zenodo with verifiable DOIs — not a single paper, but the map of the entire research program: experiments, data, findings, and the interactive tools derived from them.

Where do you want to start?

Compatible with VoiceOver and TalkBack
Take the survey (opens in a new tab)


The central question

How does mapping design influence the visual information that survives in the acoustic representation?

Initial question (ARGIRA I–III): Is Δ predictable from visual features? — The experiments showed that it is not. ARGIRA IV explains why.

A map of what
Δ does not explain

ARGIRA is not a predictive model. It is a framework for structured representational elimination: a sequence of experiments designed to determine which classes of visual representation are insufficient to capture the perceptual difference between sonification mappings.


The project evaluates three hierarchical levels of representation — low-level statistics, expanded feature spaces, and proxy semantic embeddings — using robust linear regression and Random Forest with 5-fold cross-validation. The corpus includes ~86 images from two heterogeneous visual collections: post-impressionist paintings and museum landscape photographs.

The sequence · How negative results led to the right question
ARGIRA I–III
Does the image predict Δ?
Three levels of visual representation evaluated: low-level · expanded · LLM semantic.
R² < 0 across the board
ARGIRA IV
Δ conceals two distinct mappings
OPRS and RTR modeled separately reveal opposing visual sensitivities.
OPRS R²=0.58 · RTR R²=0.25
ARGIRA V
New
Predictor inversion
edge_density dominates the visual side. hue_entropy dominates the acoustic side. n=391, 5 corpora.
ρ≈−0.10 orthogonality

Negative results are not a problem for the program — they are its mechanism. Full series · zenodo.20534327 ↗ (opens in a new tab)


Before the question,
the program.

Experiments 11–15 did not emerge in a vacuum. The project accumulated a corpus of prior observations over several months — pipelines, benchmarks, nature corpora, clustering analyses — that established the phenomena to be explained and ruled out early hypotheses. These deposits are the ground from which ARGIRA I–V grows.

Conceptual origin · Feb 2026
Computational Verification of Vogel's Model in Phyllotaxis Patterns
ARGIRA's real starting point was not sonification, but this series on the golden angle (137.508°) in sunflower seed arrangement and its application to optimal solar-panel layout. v1 uses a tool called "Argira Station" to analyze sunflower images — the same name and the same drive toward accessibility that later became the sonification pipeline. The series evolved over 10 versions: from pure geometric verification (v1–v4) to robustness against positional noise (v5–v6), advanced spatial metrics and Weyl discrepancy (v7, v9–v10), and a physical solar-shading model with real astronomical trajectories (v8). Shared function with the sonification line: both ask whether a mathematical structure (golden angle, chromatic dispersion) reliably translates into another domain (spatial, acoustic) without explicit rules.
v10 (latest) ↗ v1 (origin) ↗
Version history (v1 – v9)
v9 Same description as v10 (boundary defects fixed, Weyl discrepancy, solar-shading model v2); v10 only adds a link to an interactive results webpage. 10.5281/zenodo.18758919 ↗
v8 "Solar v2" — introduces a physically grounded photovoltaic model: real solar trajectory, hourly shadow projection, annual energy-yield estimation, and latitude sensitivity (20°N/40°N/60°N). Key finding: geometric uniformity alone is an incomplete criterion — Fibonacci retains ~98.5% of the ideal energy yield across all tested latitudes. 10.5281/zenodo.18749926 ↗
v7 Adds pair-correlation function g(r), radial density deviation (RDD), boundary-defect quantification, and Diophantine approximation. Fibonacci and the classic Vogel spiral turn out to be geometrically equivalent (U=98.86%). Halton/Sobol achieve better radial uniformity but worse global uniformity. 10.5281/zenodo.18749750 ↗
v6 Introduces Gaussian positional perturbation (σ=0.5–5 px, 100 Monte Carlo replicates) simulating real installation errors. Finding: under noise, higher-order k-nacci sequences (k=4) degrade proportionally less than Fibonacci — a robustness/optimality trade-off not visible in earlier noise-free versions. 10.5281/zenodo.18740184 ↗
v5 First systematic robustness analysis varying N∈{50,100,200,500,1000}. Fibonacci maintains U≥98% at all scales; k≥4 sequences degrade progressively from N≥500 onward. 10.5281/zenodo.18735181 ↗
v1–v4 Initial computational verification of Vogel's model (1979): the golden angle 137.508° emerges as the optimal packing solution (relative errors <10⁻⁸%, 500 synthetic instances). v1 introduces the "Argira Station" tool to analyze real sunflower images. v2–v4 shift the focus toward solar panels and add Python/NumPy simulation code. 10.5281/zenodo.18703761 ↗
Foundational correlation
Chromatic variance & spectral roughness — Replication study (N=30)
Preprint v4. Corpus of 30 canonical paintings (15 artists, 6 centuries). Main result: hue_std → roughness r = 0.687; with Shannon entropy, R² = 0.732.
10.5281/zenodo.20364540 ↗
Mapping robustness
Mapping Variants Series — Pipelines 10–13 · Scanline control (08)
Five mathematical variants (linear, log, mel, stochastic, quadratic) across 30 paintings. Result: the shape of the hue→frequency function doesn't matter; the link itself does (r = 0.878–0.951). The scanline control (pipeline 08, no hue→frequency link) gives r = −0.116 — the link is necessary.
Pipelines 10–13 ↗ Scanline 08 ↗
Mechanism
Sine Additive Pipeline — Harmonic vs. Non-Harmonic
Pipeline 09 (pure sine wave, no harmonics): r = +0.951 vs pipeline 07 (harmonic): r = +0.687. Inter-harmonic interactions obscure the effect rather than generating it.
10.5281/zenodo.20366185 ↗
Emergent spectral geometry
Emergent Corpus v1.0 — Spectral Geometry (N=14)
Separable spectral regimes without explicit harmonic rules. Clustering shows 92.9% agreement with chromatic regimes (shuffle p < 0.005). Ds acts as a structural invariant, not as a discriminative variable.
10.5281/zenodo.20393823 ↗
Generalization to nature
Nature Corpus v2 — Saturation as Predictor (N=21)
Forest and moving-landscape photographs. Saturation → harmonic count: r = 0.9644. hue_std and fractal_D are weak in this corpus (r = 0.27 and 0.03). Saturation dominates in low chromatic-variance regimes.
10.5281/zenodo.20392520 ↗
Unsupervised clustering
Clustering Dataset v1 — 74 works, k-means + PCA
k-means spontaneously isolates Rembrandt (chiaroscuro as an acoustic signature). Starry Night: visually complex, acoustically low-entropy — rhythmic complexity, not random. Signac: extreme outlier (visual PC1 = 6.41).
10.5281/zenodo.20323085 ↗
Methodological stability
Experiment 6 — RTR & Bootstrap · Experiment 7 — OPRS stability · Experiment 8 — Threshold sensitivity
Naive > OPRS > RTR hierarchy stable across 10,000 bootstraps. OPRS converges at N≈50 runs. The visual–acoustic ordering is robust within the 1000–1750 Hz cutoff range.
Exp 6 ↗ Exp 7 ↗ Exp 8 ↗
Generalization across corpora
Experiment 9 — Generalization Across Independent Corpora
Naive generalizes to independent corpora (r ≈ 0.54–0.62). RTR produces the weakest correlations. OPRS shows a corpus-dependent effect. Empirical basis for Experiments 11–15.
10.5281/zenodo.20516325 ↗
Paradox & technical note
Stereo Width: Malevich vs. Kandinsky — The inverted paradox
Malevich (near-monochromatic) produces greater stereo width (W = 0.196) than Kandinsky (W = 0.058). Interactive Web Audio API demo. Pipeline invariant confirmed in v1 and v3.
10.5281/zenodo.20353894 ↗
Formal limits of the system
Cycle Closure Experiment — 74 works, 3 epistemological layers
Can the image be reconstructed from its sonification? No — and that is a stronger result than it appears. Layer A: horizontal center of mass preserved with error < 0.05 (direct accessibility application). Layer B: significant structural correlations (local_contrast → freq r = 0.949, edge_density → effort r = 0.759). Layer C: inversion impossible — tempo ranges do not overlap by design. ARGIRA is a system for structural preservation, not inversion.
10.5281/zenodo.20322821 ↗
Spectro-temporal analysis
B_v9 — Spectro-Temporal Persistence, 72 works, regime transition
2,160 band-speed observations (6 bands × 5 speed conditions, 0.25×–1.50×). Central finding: sign inversion in the 4–8 kHz band (ρ = +0.973 → −0.022 when going from 0.25× to 0.50×). No other band inverts sign. Robustness confirmed by 4 independent tests: ordinal invariance, re-binning, jackknife, reparameterization. Independent cross-validation: hue_std predicts residual roughness with r = 0.8477.
Preprint ↗ Conference paper ↗ Dataset ↗ Aliasing v2 · r=0.8477 ↗
The intermediate pipeline · The most downloaded
Argira Pipeline v14 — N=47, r=0.9289 · 275 views · 140 downloads
Stable version prior to the expansion to v17. Includes permutation testing (N=1000), Dv and tempo experiments, batch analyzer. Negative 2D FFT result documented. The most downloaded repository in the program — empirical reference point for the canonical corpus.
10.5281/zenodo.20097327 ↗
Extended post-submission version
Argira Pipeline v16 · Dataset v17 (74 works)
Extended version following the ICAD 2026 submission (v16, N=47, r = 0.9289) and its independent expansion to the full corpus (v17, N=74, r = 0.9233). Sobel cross- validation: r = 0.9570. Negative result documented: 2D FFT with no differential signal. The actual ICAD 2026 submission was v10 (zenodo.18865019 (opens in new tab), N=45, r = 0.8869).
Pipeline v16 ↗ (opens in new tab) Dataset v17 ↗ (opens in new tab)
The base pipeline
Argira Station v15 — Standalone analyzer, 6 validated metrics
The canonical version of the pipeline: hue_std ↔ Ds r = 0.9289 (N=47, p < 0.001). Irregularity ↔ Sobel r = 0.9570. 2D FFT removed (negative result documented). Dv ↔ Ds r = 0.0229 (not significant — an honest limit of the system). The most cited repository in the entire experimental program.
10.5281/zenodo.20113356 ↗
The tool · Accessibility
Argira Sonification v23 — 4 perceptual dimensions, spatial audio
Hue → frequency · Saturation → timbre + amplitude · centroid.x → stereo pan · centroid.y → register. Correlation hue_std ↔ Ds = r = 0.9289 (N=47). Basis for the entire subsequent experimental program. Live demo at argira.eus/argira-sonification/
10.5281/zenodo.20234082 ↗
Haptic feedback · Tactile accessibility
B_v9 Dataset v2 — Haptic module · Threshold v* = k − m · hue_std
The B_v9 v2 dataset includes the haptic feedback module (client-side, no server), operational in the online analyzer. The Argira threshold law defines v* as a function of hue_std: v* = 7800 − 12500 · hue_std. For the 43 works, all v* values fall within the [4000, 8000] Hz range. Correlation hue_std ↔ Ds: r = 0.9370 (N=43). Haptics act as a perceptual channel complementary to sound for tactile exploration of artworks.
Dataset B_v9 v2 ↗ Haptic demo ↗
Spoken image description · Visual accessibility
Argira Vision v0.5 — Spoken semantic description via Claude API
Browser prototype that sends an image to the Claude API (Haiku/Sonnet) and reads the description aloud via speech synthesis. It complements ARGIRA's sonic channel with a semantic one: while ARGIRA translates visual structure into frequency and timbre, Vision translates visual semantics into language. It is not real-time navigation — it is auditory exploration of visual content. In development · experimental
10.5281/zenodo.18989327 (v0.5) ↗

Additional corpora: Unified dataset of 74 images · zenodo.20402536 ↗ · Space A=f(H,S,I) · zenodo.20388416 ↗ · Dataset B_v9 (43 works, threshold v*) · zenodo.20357570 ↗

The correlation that
started it all

Analyzing 45 works with Argira's conversion rule, a statistically robust correlation emerged between a painting's chromatic variability and the sonic complexity of its sonification.


This correlation —r = 0.8869, R² = 0.8417, p < 0.001— was not designed. It emerged from the data. It is the reason the project exists: color appears to be related to sound.

Pearson r = 0.8869 · R² = 0.8417 · p < 0.001 · N = 45 · Argira v3.5
hue_std (variabilidad de tono) → Ds (dimensión espectral / complejidad sonora)
Malevich Kandinsky r = 0.8869 hue_std → Ds →

16 works shown · Correlation computed on N = 45 works (full dataset) · Zenodo v10 (opens in new tab)


This correlation is the starting point, not the conclusion. The question it opens is: what visual features explain this link? Experiments 11–15 attempt to answer it — and find that standard visual representations are not sufficient.


The sonic museum argira.eus/argira-sonification/ lets you hear this correlation in action: the 16 works ordered from lowest to highest chromaticity, from Malevich to Kandinsky.



Three levels.
The same result.

Study summary

~86
Initial corpus
47–72
Final valid sample
11–15
Experiments
3
Models evaluated
Ridge · Huber · RF
5-fold
Cross-validation
0.10
Best R² CV
Does not generalize

From perceptual divergence
to structural asymmetry

ARGIRA IV re-examines the negative results of the previous phases from a different perspective. Instead of directly modeling Δ as the target variable, it separately analyzes the two mappings that compose it: OPRS and RTR. This methodological shift reveals a previously hidden structural asymmetry.

Independent mapping modeling — ARGIRA IV

O
Mapping A
OPRS
Ridge Regression
0.582
R² CV
Random Forest
0.528
R² CV
p ≈ 0.002 · permutation test

Preserves a substantial portion of the image's visual structure, particularly that related to spatial texture and roughness.

R
Mapping B
RTR
Ridge Regression
0.247
R² CV
Random Forest
0.086
R² CV
Weak relationship · No significant permutation signal

Much weaker relationship with spatial structure. It appears to respond mainly to global chromatic features.

Main finding · ARGIRA IV

The negative results of ARGIRA I–III do not indicate an absence of image–sound relationships.

The difficulty arises from directly modeling a composite variable (Δ) that combines two transformations with distinct and partially opposing sensitivities. When both mappings are studied separately, clearly differentiated predictive structures emerge.

ARGIRA IV · zenodo.20531660 (opens in new tab)


Predictor inversion
across domains

ARGIRA V investigates whether the same visual variables predict roughness across visual and acoustic domains. The study combines five independent visual corpora (n=319 images) and one acoustic corpus (n=72 sonifications), with 391 cases analyzed in total.

Visual domain
edge_density
ρ = 0.49–0.84
Dominant predictor stable across corpora A–D
lum_contrast: secondary (ρ = 0.56–0.74)
Inversion
Acoustic domain
hue_entropy
ρ = 0.595
Dominant predictor in the acoustic corpus (n=72)
edge_density here: secondary (ρ = 0.428)
Total corpus
391
5 visual corpora + 1 acoustic
Orthogonality
ρ ≈ −0.10
edge_density ↔ hue_entropy (nearly orthogonal)
Regimes
3
Chromatic · Grayscale-collapse · High-edge
DOI
10.5281/zenodo.
20534328 ↗
Architectural implication

The results support a multilayer sonification architecture where structural information (edge_density) and chromatic information (hue_entropy) must be represented independently, since they contribute non-redundant information to distinct roughness domains.


Structural convergence
in predictor space

ARGIRA VI expands the canonical corpus from 187 to 227 images (96 canonical paintings by 19 identified artists, plus 131 synthetic controls), using the new Sonification Pipeline v3.4 with a three-channel orthogonal architecture: hue_entropy controls the base frequency, edge_density + fractal_D control modulation depth (predictor derived by OLS, R² = 0.640), and luminance_contrast controls decay time.

Expanded corpus
227
96 canonical (19 artists) + 131 controls
Channel B (OLS)
R² = 0.640
edge_density + fractal_D → mod_depth
Structural convergence
80 pairs
Δedge>0.15 with near-identical mod_depth

The central finding is that visually very different paintings can converge toward nearly identical acoustic states within ARGIRA's mapping architecture: 80 pairs of canonical works satisfy |Δfractal_D| < 0.005 with |Δedge_density| > 0.15, yet still produce practically equal modulation depth values. At the artist level, recognizable structural signatures emerge in the fractal_D × edge_density space —Turner consistently occupies the high-fractal/low-edge region, Caravaggio the low-fractal/mid-edge region— without using any explicit stylistic label.

DOI: 10.5281/zenodo.20553976 · Pipeline v3.4 (companion): 10.5281/zenodo.20550714


A matrix to
measure association, not assume it

ARGIRA VII introduces the Visual-Acoustic Association Matrix (VAAM), a systematic representation of how five primary visual descriptors (hue_std, hue_entropy_bits, edge_density, fractal_D, luminance_contrast) and three interaction terms associate with seven acoustic dimensions, over a canonical corpus of 117 paintings processed with a single sonification operator (M1).

Corpus
N = 117
+ 110 synthetic controls
fractal_D
S ≈ 0.27
acoustic inertia — minimal predictive power
Imposed structure
S ≈ 1.000
hue_std→odd_bias · lum_contrast→decay_s

Four findings organize the study. First, the visual predictors group into two functionally independent families: a chromatic family (hue_std, hue_entropy_bits) that maps preferentially to frequency-related dimensions, and a structural family (edge_density, luminance_contrast) that maps to temporal and texture dimensions. Second, hue_std and hue_entropy_bits are non-redundant observables —partial correlation shows they independently predict distinct acoustic targets. Third, fractal_D shows minimal predictive power across all acoustic dimensions (maximum S ≈ 0.27), which contrasts with its documented relevance in empirical aesthetics and suggests that this sonification architecture transmits fractal information poorly. Fourth, the roughness dimension concentrates the matrix's highest nonlinearity indices.

Central methodological distinction

Two associations reach near-perfect values (S ≈ 1.000) —hue_std→odd_bias and luminance_contrast→decay_s— but are not interpreted as empirical discoveries: they arise directly from deterministic equations implemented in the sonification operator itself, and are explicitly reported as structure imposed by the operator, not as emergent properties of the corpus. Distinguishing imposed relationships from observed associations is the central methodological contribution of this repository. The question of whether visual-acoustic associations survive across structurally distinct sonification operators remains open and is the objective of ARGIRA VIII (opens in new tab). The answer: hue_std→odd_bias does not survive — it is falsified under a full architectural reorganization (M3b) and reclassified as an operator artifact, not an invariant property of the corpus.

DOI: 10.5281/zenodo.20554745


The conceptual turn · ARGIRA V

The question is no longer what sound does an image produce?
But rather: what aspect of the image do I want the user to perceive through sound?

A mapping is a
perceptual hypothesis

ARGIRA IV reveals that no sonification algorithm translates an image: it selects which visual dimension deserves to survive in the sound. This selection is not technical, it is epistemological. The mapping's designer decides —implicitly— what the image is for the listener.


A direct consequence follows from this principle: there is no single correct sonification of an image. There are as many valid sonifications as there are visual properties considered relevant to communicate.

ARGIRA IV · zenodo.20531660 (opens in new tab)

Space of possible mappings · Theoretically infinite
Table of possible mappings: visual property, acoustic mapping, and resulting perceptual experience
Visual property Possible mapping What the user perceives
Color (Hue) hue → frequency Chromatic climate
Saturation saturation → amplitude Emotional intensity
Brightness luminance → pitch Lightness / darkness
Edges edge density → roughness Geometry
Texture roughness → modulation Relief
Symmetry symmetry → consonance Order
Entropy entropy → spectral noise Complexity
Contrast contrast → dynamic range Tension
Depth depth → reverberation Space
Motion optical flow → tempo Dynamism

ARGIRA currently explores: hue, chromatic entropy, spatial roughness.

Four types of
perceptual filter

Building on the results of ARGIRA IV, it is possible to sketch a first classification of mappings according to the type of visual property they prioritize and the perceptual experience they generate in the listener.

ARGIRA IV · zenodo.20531660 (opens in new tab)

Type I
Structural
Preserve shape

Convey the geometry of the image: edges, contours, orientation, composition. The listener can "feel" the visual structure.

Examples:
OPRS · edge sonification · contour tracking
Goal: feel the geometry
Type II
Chromatic
Preserve color

Translate the image's palette into acoustic dimensions. Hue, saturation, and brightness as three independent channels of chromatic information.

Examples:
RTR · hue→frequency · color harmony
Goal: feel the palette
Type III
Topological
Preserve spatial organization

Map the position within the image to acoustic space. X position to panning, Y position to pitch. The listener navigates the image like a territory.

Examples:
Argira's "Play" mode · spatial audio scanning
Goal: navigate the image
Type IV
Expressive
Preserve emotion or atmosphere

Convey the emotional or expressive character of the image — not a measurable property, but the feeling it generates.

Examples:
Pending empirical definition
Goal: feel the character

An image as an orchestra:
each property, a voice.

Most sonification systems apply a single, sequential mapping. The multilayer architecture proposes something different: each visual property generates an independent, simultaneous acoustic layer, like the sections of an orchestra.

SOURCE IMAGE
COLOR
→ Channel 1
TEXTURE
→ Channel 2
GEOMETRY
→ Channel 3
DEPTH
→ Channel 4
↓ Final acoustic mix

Δ cannot be modeled robustly using the visual representations evaluated.

Increasing representational complexity —from basic statistics to expanded spaces to semantic embeddings— does not consistently improve the prediction of Δ. ARGIRA IV suggests this limitation arises because Δ combines two mappings with different visual sensitivities, hiding predictive structures that emerge when each mapping is analyzed separately.

ARGIRA II · zenodo.20526610 (opens in new tab)
ARGIRA III · zenodo.20530682 (opens in new tab)

Best cross-validated R² by visual representation level
Level Best CV R² Status
Low level
Exp 11–13
−0.44
No signal
Mid level
Exp 14
+0.10
Marginal
Semantic
Exp 15
−0.337
No signal

Sonification algorithms
are not neutral translators

Each mapping selects, preserves, and discards different properties of the image. The choice of mapping determines which aspects of the work survive in the acoustic representation.

01
Sonification system design
The choice of mapping is not neutral: it determines which visual structure is conveyed acoustically.
02
Visual accessibility
Direct implications for designing tools for people with visual impairments.
03
Multimodal interfaces
Empirical foundation for designing image-to-sound systems with perceptual fidelity.
04
Museums and auditory exploration
Tools for the auditory exploration of visual works in museum contexts.

Central finding of the ARGIRA project

Sonification algorithms are not neutral translators of visual information.

Each mapping selects, preserves, and discards different properties of the image, generating distinct acoustic representations even when operating on the same visual data.

The negative results did not indicate the absence of image-sound relationships.
They indicated the existence of multiple sensory translation mechanisms hidden beneath a single composite measure.

From studying a correlation to proposing a theory of mapping design.

If this result is consolidated with more mappings, ARGIRA could move from analyzing Δ as a statistical variable to proposing a theoretical framework for mapping design in multimodal accessibility — a considerably broader contribution than explaining a perceptual difference between two algorithms.

What the brain receives

Not "the sound of the image."

An acoustic interpretation of a specific property of the image.

OPRS → structure · roughness · texture
RTR → atmosphere · color · chromatic distribution

Several ways to tell
the same project

ARGIRA is an auditable instrument, not a landing page. There is material dedicated to explaining it at different levels of depth — from a single sentence for someone new to the topic, to the formal architecture for an ICAD reviewer. All of them describe the same central idea: moving from a black box to an auditable process. This page (the publication index) is just one of those ways.

Questions and answers · Glossary
Knowledge Assistant de ARGIRA / RMA
Answers specific questions ("What is ARGIRA?", "What is RMA?", "What is hue_std?", "What does the correlation show?") always citing the official documentation — with no language model and no internet connection. Includes a technical glossary (pipeline features: hue_std, edge_density, fractal_D...) and a conceptual glossary (Feature, Target, Bootstrap CI, Dependency Network...), with a confidence indicator on every answer (✓ Confirmed, ≈ Partial, ? Unconfirmed). Recommended entry point for anyone arriving without prior context or needing to verify an exact term.
Open the Assistant (opens in new tab)
Guided explainer · 12 min
Explainer ARGIRA ↔ RMA
Explains the question that started the whole project (what information from a painting survives its conversion into sound), what RMA means in simple terms (Y = g(X) + R — the residual R is what the model doesn't explain), and how the Knowledge Explorer brings together pipeline, data, evidence, hypothesis, and analysis in a single instrument. Offers several reading lengths (2 min, 10 min, technical) depending on the time available.
Open the Explainer (opens in new tab)
Academic material · ICAD 2026
ICAD poster in four formats
"Crossmodal Correspondences Emerging from Painting Sonification" — presented at ICAD 2026 (31st International Conference on Auditory Display, ESMUC). Available as a printable A1 PDF, a responsive HTML twin of the poster, an interactive desktop version with the full 74-work corpus (scatter map, tooltips, correlations), and a mobile companion designed to scan the printed poster's QR code during the session.
Open the mobile companion (opens in new tab)
Interactive experience · Accessible
Museo Sonoro Accesible
The 16 works converted into sound, navigable with touch gestures: Tocar (position, color, and frequency in Hz for each zone), Barrido (continuous exploration), Tour espacial (automatic zone-by-zone description), and Braille in development. Includes Lupa temporal to slow audio down to 0.5× and perceive chromatic texture, a Mapa Perceptual where the sound changes in real time based on cursor or finger position, a Test Auditivo of crossmodal perception ("can you hear the color?"), the option to upload your own image to generate its acoustic signature, an Asistente with a technical and conceptual glossary, and explainer videos. Designed with VoiceOver and TalkBack support.
Each work exposes its perceptual metrics along five axes (Color, Structure, Contrast, Complexity, Spatial organization) and its raw values: hue_std, edge_density, luminance_contrast, fractal_D, stereo_pan, sat_mean, and centroid (x, y) — the same variables used in the r = 0.8869 correlation.
Open the Museo Sonoro (opens in new tab)

This page (the complete index of publications and audits, with verifiable DOIs) is one more way: the most exhaustive, designed for anyone who needs complete data-by-data traceability.

The ARGIRA program
on one page

A sonification research project that systematically evaluates which classes of visual representation can — and cannot — predict the perceptual difference between acoustic mappings. Five studies, all open access on Zenodo.

Research question

How does mapping design influence the visual information that survives in the acoustic representation? Which visual features predict the perceptual difference Δ = OPRS − RTR?

Corpus and method

~86 images (post-impressionist paintings + museum photographs). Ridge Regression and Random Forest with 5-fold CV. Three levels of visual representation evaluated hierarchically. All materials reproducible on Zenodo.

ARGIRA program: five studies with research question, cross-validated R² result, and DOI link
Study Question Result DOI
ARGIRA I Low-level features (8 variables) — Exp 11 R² CV < 0 20524644 ↗ (opens in new tab)
ARGIRA II PCA-expanded features (209 variables) — Exp 11–15 R² CV ≈ 0.10 20526610 ↗ (opens in new tab)
ARGIRA III LLM semantic descriptors — Claude Haiku, n=72 R² CV = −0.162 20530682 ↗ (opens in new tab)
ARGIRA IV OPRS and RTR modeled separately — structural asymmetry OPRS R²=0.582
RTR R²=0.086
20531660 ↗ (opens in new tab)
ARGIRA V Cross-corpus predictor inversion — n=391, 5 visual corpora + 1 acoustic edge_density ≠ hue_entropy
ρ≈−0.10
20534328 ↗ (opens in new tab)
Central finding

Δ cannot be robustly modeled using the visual representations evaluated. ARGIRA IV reveals why: OPRS and RTR preserve structurally distinct visual properties. ARGIRA V confirms that the asymmetry generalizes across domains and corpora.

Cite as

Ranero García, J. (2026). ARGIRA: Sonification and Perceptual Metrics (Complete series I–V). Zenodo.
doi.org/10.5281/zenodo.20534327 ↗ (opens in new tab)

ICAD 2026 · Poster #4302 · Barcelona · Jul 2026


Open access
to all materials

ARGIRA Series I–XII

Correlation → structural asymmetry → invariance → falsification
XII
XII — Transition Matrix and Structural Observer Framework
Dataset · Zenodo · 6 Jun 2026 · Transition matrix · structural observers · 7 figures

Formalizes a "structural observers" framework using a transition matrix, closing the series with a methodological synthesis of the entire ARGIRA program.

10.5281/zenodo.20569073
XI-A
ARGIRA XI-A — GLCM Residual Structure Analysis of ARGIRA X
Dataset · Zenodo · 6 Jun 2026 · N=227 · glcm_global_contrast · ΔR² ≈ 0.064

Applies GLCM texture analysis to the unexplained residuals of ARGIRA X (N=227), to check whether visual structure remains uncaptured. Finds a modest fit improvement (ΔR² ≈ 0.064).

10.5281/zenodo.20568815
X
ARGIRA X — Empirical Spectral Analysis of Operator-Dependent Cross-Modal Coupling in a Visual-to-Acoustic System
Working paper · Zenodo · 5 Jun 2026 · SVD + Procrustes · rank-1 (M1) → rank-3 (M3)

Empirical spectral analysis (SVD + Procrustes) of how visual-acoustic coupling changes depending on the sonification operator used, comparing rank-1 (M1) to rank-3 (M3) configurations.

10.5281/zenodo.20556440
IX
ARGIRA IX: Channel Covariance and Operator-Induced Migration in Visual-to-Acoustic Perceptual Decomposition
Dataset · Zenodo · 5 Jun 2026 · tempo_color/fractal/space decomposition

Analyzes covariance between perceptual channels (tempo_color, fractal, space) and how that covariance reorganizes — "migrates" — when the sonification operator used is changed.

10.5281/zenodo.20555763
VIII
ARGIRA VIII: Invariant Survival Analysis of Visual–Acoustic Associations Across Competing Sonification Operators
Dataset · Zenodo · 5 Jun 2026 · N=117 · falsification of hue_std→odd_bias · dep_arquitectura=21

Answers the question left open by ARGIRA VII: does the hue_std→odd_bias association survive a complete change of sonification architecture (N=117, M3b reorganization)? It does not survive — it is reclassified as an operator artifact rather than an invariant property of the corpus.

10.5281/zenodo.20555305
VII
ARGIRA VII: Visual–Acoustic Association Matrix (VAAM) for a Canonical Painting Corpus
Dataset · Zenodo · 5 Jun 2026 · N=117 paintings + 110 controls · fractal_D S≈0.27

Introduces the Visual-Acoustic Association Matrix (VAAM), cross-referencing 5 visual descriptors with 7 acoustic dimensions over 117 paintings + 110 synthetic controls. The chromatic and structural families predict distinct, non-redundant acoustic dimensions; fractal_D has almost no predictive power. Explicitly distinguishes associations imposed by the sonification operator from real properties of the corpus.

10.5281/zenodo.20554745
VI
ARGIRA VI: Emergent Structure in Visual-to-Acoustic Mapping — Canonical Corpus of 96 Paintings
Software · Zenodo · 5 Jun 2026 · Corpus N=227 · 19 artists · structural convergence (80 pairs)

Extends the canonical corpus to 227 works by 19 artists. Shows that visually very different paintings can converge on nearly identical acoustic states (80 pairs with large edge difference but nearly equal modulation depth), and that per-artist structural signatures emerge without using any style label.

10.5281/zenodo.20553976
V
ARGIRA V: Cross-Corpus Evidence for Predictor Inversion in Visual-to-Acoustic Mapping
Dataset · Zenodo · 4 Jun 2026 · n=391 (5 visual corpora + acoustic)

Across 391 cases (5 visual corpora + 1 acoustic), finds a predictor inversion across domains: edge_density dominates visually (ρ 0.49–0.84), but hue_entropy dominates acoustically (ρ = 0.595). Supports a sonification architecture with independent chromatic and structural channels.

10.5281/zenodo.20534328
IV
ARGIRA (IV): Structural Asymmetry in Sonification Mappings
Software · Zenodo · 3 Jun 2026 · OPRS vs RTR · Independent modeling

Turning point of the series: instead of modeling Δ as a combined variable, it separates the two mappings (OPRS and RTR) and models them independently. OPRS retains predictive structure (R² CV ≈ 0.58, tied to texture and roughness); RTR barely retains it (R² CV ≈ 0.25, tied to overall color). This explains why I–III found no signal: Δ was mixing two distinct visual sensitivities.

10.5281/zenodo.20531660
III
ARGIRA (III): Systematic Representational Failure — LLM Semantic Descriptors Cannot Predict Perceptual Sonification Divergence
Dataset · Zenodo · 3 Jun 2026 · Claude Haiku · n=72 · R² CV = −0.162

Replaces numerical features with semantic descriptors generated by an LLM (Claude Haiku) — the "highest-level" visual representation evaluated in the series. The predictive signal disappears entirely (R² CV = −0.337 in the associated experiment 15), reinforcing that the problem was not a lack of representational complexity.

10.5281/zenodo.20530682
II
From Correlation to Failure: Limits of Visual Feature Spaces in Predicting Perceptual Sonification Differences
Experiments 11–15 · Preprint · Zenodo · 2026

Repeats the ARGIRA I test with a much larger feature space (~209: HSV histograms, Laws filters, per-quadrant statistics, Sobel) reduced with PCA. The signal improves but remains marginal and unstable (R² CV ≈ 0.10) — it does not generalize.

10.5281/zenodo.20526610
I
Visual Predictors of Acoustic Perceptual Distance in Image Sonification
Experiment 11 · Software · Zenodo · 2026

First experiment in the series: tests whether ~8 low-level visual features (hue, saturation, edge density, roughness, luminance contrast) predict the perceptual divergence Δ between two sonification mappings. No signal (R² CV < 0).

10.5281/zenodo.20524644
Prior corpus · Experiments 1–10
E10
Chromatic variance and spectral roughness in image sonification — Replication study (N=30)
Preprint v4 · Zenodo · 24 May 2026 · 30 paintings · hue_std → roughness r = 0.687 · R² = 0.732 with entropy

Replicates on a larger corpus (N=30, vs N=9–10 in v1–v3) the hue_std→spectral roughness correlation found by chance while comparing the original mono pipeline with the stereo version (Argira v23). The correlation drops from r≈0.85 to r=0.687 as the sample grows, but remains significant (p<0.001); combined with entropy it rises to R²=0.732.

10.5281/zenodo.20364540
Version history (v1 – v3)
v3 Introduces an independent "naive" harmonic pipeline (without fractal metrics, stereo spatialization, MFCC, or adaptive weighting) to check whether the correlation survives outside the Argira architecture. With N=9, r≈0.86 — it remains high despite the radical simplification, and the Vermeer outlier is preserved in the independent pipeline. 10.5281/zenodo.20361230 ↗
v2 Cross-validation addendum: argues that this emergent (undesigned) correlation refutes the circularity objection raised against Argira's main correlation (hue_std → Ds, by design). Same N=10 and r=0.8477 as v1; adds the methodological argument as a PDF. 10.5281/zenodo.20252194 ↗
v1 Initial unsought finding: comparing the museum's original mono pipeline with the Argira v23 stereo implementation, hue_std predicts residual spectral roughness with r=0.8477 (N=10) — a correlation emerging from the pipeline's mathematical structure, not coded as a target parameter. 10.5281/zenodo.20245918 ↗

Note: when the sample is expanded from N=9–10 to N=30 in v4, the correlation drops from r≈0.85 to r=0.687 (though it remains significant, p<0.001) — a digitization artifact (Black Square) is also documented, which, when excluded, raises r to 0.790.

E9
Experiment 9 — Generalization of Hue Dispersion–Roughness Across Independent Corpora
Dataset · Zenodo · Jun 2026 · Naive r ≈ 0.54–0.62 · Naive > OPRS > RTR

Checks whether the relationship between hue dispersion and roughness generalizes to independent corpora, comparing three pipelines (naive, OPRS, RTR). The naive pipeline generalizes best (r ≈ 0.54–0.62), above OPRS and RTR.

10.5281/zenodo.20516325
E8
Experiment 8 — Robustness of Visual–Acoustic Correspondence Across Roughness Thresholds
Report · Zenodo · 2 Jun 2026 · Stable hierarchy at 1000–1750 Hz · N=69

Tests the robustness of the visual-acoustic correspondence by varying the roughness threshold across a corpus of 69 works. The result hierarchy remains stable in the 1000–1750 Hz range.

10.5281/zenodo.20513550
E7
Experiment 7 — Stability of OPRS Under Repeated Random Sampling
Report · Zenodo · 2 Jun 2026 · Practical convergence at N_RUNS ≈ 50–100

Evaluates whether OPRS produces stable results when random sampling is repeated. It converges in practice around 50–100 repetitions.

10.5281/zenodo.20513415
E6
Experiment 6 — RTR, Bootstrap Analysis and the Role of Relative Order
Report · Zenodo · 2 Jun 2026 · 10,000 bootstraps · RTR r = −0.417 · relative order matters

Analyzes RTR with 10,000 bootstrap resamples; finds r = −0.417 and confirms that the relative order of values, not just their magnitude, influences the result.

10.5281/zenodo.20513340
EC
Emergent Corpus v1.0 — Spectral Geometry from Minimal Hue→Frequency Sonification
Dataset · Zenodo · May 2026 · N=14 · clustering 92.9% · Ds as invariant

Corpus of 14 works sonified with the minimal hue→frequency rule. Unsupervised clustering correctly groups 92.9% of cases, showing that the spectral dimension (Ds) acts as an invariant even with this simplified version of the pipeline.

10.5281/zenodo.20393823
NC
Nature Corpus v2 — Saturation as Predictor of Emergent Harmonic Count (N=21)
Dataset · Zenodo · May 2026 · Mobile photographs · saturation → count r = 0.9644

21 landscape photographs taken with a mobile phone. Color saturation predicts the number of emergent harmonics with a very high correlation (r = 0.9644).

10.5281/zenodo.20392520
CL
Clustering Dataset v1 — Unsupervised Sonic Grouping Across 74 Artworks
Dataset · Zenodo · May 2026 · k-means k=4 · PCA · Rembrandt cluster · Signac outlier

Unsupervised clustering (k-means, k=4, with PCA) over 74 sonified works. Identifies a cluster dominated by Rembrandt and Signac as a corpus outlier.

10.5281/zenodo.20323085

RMA Line — Structural Loss Audit

What survives a deterministic transformation, domain by domain
R↔A
ARGIRA ↔ RMA v1.4 — Reproducible Technical Artifacts for Auditing Candidate Structural Information Loss in Cross-Modal Translation
Software · Zenodo · 5 Jul 2026 · Expansion Score: RF vs. interaction model · last version before the Knowledge Explorer

Last version of the RMA line in the static package format (documents, reports, scripts) — not the final version of the RMA line overall, which is the Knowledge Explorer v1.8 that supersedes it. Adds the Expansion Score audit, comparing Random Forest against a low-order interaction model.

10.5281/zenodo.21207188
Version history (v1 – v1.4)
v1.3 Closes tempo_bpm, the last target shared by three candidates (exact recovery with a nested per-channel model). Completes Layer 2 of Candidate 4 (correlation, MI, PCA, feature importance) and adds bootstrap stability analysis: reveals that, of the confounder pairs reported as ~0 in Layer 2, only one turns out to be stable under resampling — the rest is indistinguishable from sampling noise. Does not reopen any previously closed result; adds a stability measure that point estimates alone could not provide. 10.5281/zenodo.21154398 ↗
v1.2 Mandatory Dependency Mapping (Step 2.0) + 2 new candidates (Texture, Color) + Global Structural Representation. 10.5281/zenodo.21144059 ↗
v1.1 Incorporates LOO-CV + bootstrap CI on Design A2 (Spatial Arrangement). 10.5281/zenodo.21132235 ↗
v1 First version. Design A (circular, control) and Design A2 (non-circular) for Spatial Arrangement. 10.5281/zenodo.21131138 ↗
KEx
ARGIRA ↔ RMA Knowledge Explorer v1.8 — Interactive Dependency Network and Residual Model Auditing Framework
Software · Zenodo · 16 Jul 2026 · Offline explorer · Evidence/Research/Analysis · Pyodide in-browser

Current and most recent version of the entire RMA line. Replaces the static package of documents and scripts with an interactive application that runs offline in the browser (Python engine via Pyodide), organized into three layers: Evidence (audited results), Research (exploratory hypotheses), and Analysis (on-demand exploration, whose results stay local to the browser session and are never automatically promoted to Evidence). v1.8.0 is a stabilization release focused on release consistency, artifact integrity, and documentation clarity — it integrates the confirmed standalone HTML build as the canonical artifact, verifies that v1.7 functionality is preserved (including the v3.4 and v3.5.7 knowledge representations) and the Research Mode separation, and updates documentation and citation metadata; there are no changes to RMA's computational methodology, the Bootstrap CI calculation, the residual R² analysis, or the statistical classification logic. It documents as an unconfirmed edge case a possible message-persistence scenario in the Bootstrap interface, reviewed via static code inspection: it could only occur under a very specific interaction-timing condition with controls locked during the transition between Bootstrap completion and interface unlock; it has not been reproduced in normal use and no code change was applied in v1.8.0.

10.5281/zenodo.21400719
Version history (v1.4 – v1.7)
v1.7 Integrates the ARGIRA v3.5.7 pipeline representation (versus v3.4 in previous versions), with the sonificacion_resultados_108.csv dataset (108 paintings) and additional hue descriptors. RMA engine unchanged; only the represented feature space is expanded. 10.5281/zenodo.21384719 ↗
v1.6 Embeds the full corpus (N=227) directly in the HTML, allowing Analysis Mode to work without an external CSV. Renames the main file to ARGIRA_RMA_Knowledge_Explorer.html. Adds report export in HTML/JSON/PDF and graph export in SVG. Note from the entry itself: from this version onward, the license changes from dual MIT/CC BY-NC-SA to fully CC BY-NC-SA 4.0. 10.5281/zenodo.21258130 ↗
v1.5 Changes the main artifact format: replaces the static package of documents/scripts (used through v1.4) with a single interactive web application that runs offline in the browser, with the rma_core.py engine running via Pyodide. Introduces the explicit separation into three layers — Evidence (audited, read-only results), Research (exploratory hypotheses), and Analysis (on-demand execution, never automatically promoted to Evidence). No changes to the frozen pipeline or to previous results. 10.5281/zenodo.21229821 ↗
v1.4 Last version of the ARGIRA↔RMA line in the static package format (documents, reports, scripts). Adds the Expansion Score audit (RF vs. low-order interaction model). From this point the line transitions to the Knowledge Explorer format. 10.5281/zenodo.21207188 ↗
MIDI
Residual Mapping Auditor (RMA) — Confirmatory Feature-Validation Pipeline for Symbolic MIDI Features (MAESTRO / Aria-MIDI)
Software · Zenodo · 5 Jul 2026 · duration_std confirms on MAESTRO (R²=0.32) but not on Aria-MIDI (R²=−0.26)

Applies the RMA auditor to symbolic MIDI score features. The duration_std variable is confirmed on the MAESTRO corpus (R²=0.32) but does not hold on Aria-MIDI (R²=−0.26).

10.5281/zenodo.21202724
OCR
RMA-OCR v1 — Auditing Structural Information Preservation Under Optical Character Recognition
Software · Zenodo · 4 Jul 2026 · FUNSD corpus · Tesseract 5.3.4 · 20 structural variables, 4 phases

Audits how much structure survives when documents are processed through optical character recognition (Tesseract 5.3.4), on the FUNSD corpus, with 20 structural variables evaluated across 4 phases.

10.5281/zenodo.21193539
JPEG
RMA-JPEG v1 — Auditing Structural Information Preservation Under JPEG Compression
Software · Zenodo · 3 Jul 2026 · N=109 paintings · QF=10 · color r≥0.98 · edge_density r=0.77

Audits structural loss when compressing 109 paintings to JPEG (quality QF=10). Color is preserved almost intact (r≥0.98) but edge density degrades more (r=0.77).

10.5281/zenodo.21157733
RMA3
Residual Mapping Audit (RMA) v3 — A Framework for Model-Class Dependent Residual Analysis in Parametric Systems
Software · Zenodo · 1 Jul 2026 · Permutation + invariance + LOO-CV/bootstrap · 169/169 tests

Formalizes the general audit framework: permutation test, invariance analysis, and leave-one-out/bootstrap cross-validation, with 169 of 169 tests passed. v2 fixed three reproducibility bugs and applied it both to painting sonification (ARGIRA, N=227) and to an unrelated energy-efficiency dataset (UCI, N=768), with qualitatively different results across domains.

10.5281/zenodo.21114015
Version history (v1 – v2)
v2 Fixes three bugs that affected the reproducibility and benchmark results of v1; all empirical results were recalculated with the corrected code. Applied to painting sonification (ARGIRA, N=227) and UCI Energy Efficiency (N=768), with qualitatively different results across domains and across targets within the same dataset. 10.5281/zenodo.21048838 ↗
RS
Residual Structure in Painting Sonification: Correlation Patterns in Chromatic and Acoustic Features
Preprint · Zenodo · 26 Jun 2026 · hue_std→roughness r=0.923 (N=74) · phyllotaxis · H1–H3

Documents the hue_std→roughness correlation across a corpus of 74 works (r=0.923). The entry also labels "phyllotaxis" and three hypotheses (H1–H3) without developing them further — see the DOI for the full content.

10.5281/zenodo.20925302

Accessibility · Visual characterization

Alt-text with communicated uncertainty, more recent than the RMA line
Lex
ARGIRA v2 Lexica Subcorpus — Visual Style Annotation and Heuristic Characterization (v1.1.0) New
Dataset · Zenodo · Aug 7, 2026 · N=23 · LEXICA A–C.2 · the human/AI boundary does not generalize as an origin detector

Evidence package documenting the LEXICA experimental sequence (A through C.2) on an exploratory subcorpus of 23 images from Lexica, within ARGIRA's deterministic visual characterization heuristic. LEXICA-A: human annotation of visual style with a closed taxonomy defined before running ARGIRA (photorealistic, painterly, flat illustration, digital composition), frozen prior to any heuristic analysis. LEXICA-B: unmodified run of the ARGIRA heuristic, providing a descriptive comparison between human-assigned style and ARGIRA's outputs (predictions, confidence values, margins). LEXICA-C.1: exploratory analysis of the disagreements between human descriptions and heuristic outputs, treated as two distinct representations of visual structure rather than as classification errors. LEXICA-C.2 (added in this version 1.1.0): projects the 23 Lexica images onto a decision boundary created in an earlier, independent experiment (Fase A, which compared 107 human paintings with 20 Lexica images from a visually homogeneous pictorial collection); the boundary was not retrained or validated with these 23 images, only used as a projection reference. The result suggests that the earlier separation does not generalize as an image-provenance detector (human vs. AI origin), but instead appears related to the visual characteristics of the original pictorial collection. Includes a post-hoc coherence audit between LEXICA-C.1 and LEXICA-C.2 that does not modify any prior result. This is an exploratory evidence package: it is not a trained classifier, not a benchmark dataset for AI detection, and not a statistical validation of ARGIRA's performance — the subgroups are small and the results should be interpreted as characterization of heuristic behavior, not as an estimate of general performance. Version 1.1.0 adds LEXICA-C.2 and the Section 8 coherence audit relative to the earlier v1.0.0; no file content from that earlier version was modified.

10.5281/zenodo.21836533

The full accessibility line (ARGIRA v1.4.x, fractal_D validations, synthetic RMA) is documented in the Accessibility ↓ section.

Historical material · Pipelines and variants

Development prior to the numbered series, kept for traceability
Show 21 prior development deposits (v14–v21, mapping pipelines, B_v9 dataset) · full series up to v23 featured above
WAV
Argira Station — Sonified Paintings Audio Dataset (N=45 WAV files)
Dataset · Zenodo · Mar 7, 2026 · Audio output from pipeline v3.5 · hue_std→frequency, Dv→granular texture, tempo→spatial distribution

45 audio files generated by pipeline v3.5, with hue_std mapped to frequency, Dv to granular texture, and tempo to spatial distribution.

10.5281/zenodo.18899967
MV
Mapping Variants Series — Pipelines 10–13 (log, mel, stochastic, quadratic)
Software · Zenodo · May 24, 2026 · r = 0.878–0.951 · scanline control r = −0.116

Compares four hue→frequency mapping functions (log, mel, stochastic, quadratic), with correlations between r=0.878 and r=0.951, against a scanline control with r=−0.116.

10.5281/zenodo.20366726
SA
Sine Additive Pipeline — Harmonic vs. Non-Harmonic Comparison
Software · Zenodo · May 24, 2026 · Pipeline 09 r = +0.951 · Pipeline 07 r = +0.687

Compares a harmonic pipeline (09, r=+0.951) with a non-harmonic one (07, r=+0.687).

10.5281/zenodo.20366185
P09
Sine additive sonification pipeline — Comparison of 3 architectures (07/09/08)
Software · Zenodo · May 24, 2026 · 07 naive r=+0.687 · 09 sine r=+0.951 · 08 scanline r=−0.116 · N=30

Compares the naive pipeline (07, r=+0.687), harmonic sine (09, r=+0.951), and control scanline (08, r=−0.116), with N=30.

10.5281/zenodo.20366184
P08
Argira Scanline Pipeline v2 — Spatial mapping without hue→frequency (control)
Software · Zenodo · May 24, 2026 · N=26 · r=−0.116 (p=0.572) vs naive r=+0.710 (p<0.001) · corrects FREQ_MAX 2000→10000 Hz

Control pipeline with no hue→frequency mapping; corrects FREQ_MAX from 2000 to 10000 Hz. N=26, r=−0.116 (p=0.572) against the naive pipeline's r=+0.710 (p<0.001).

10.5281/zenodo.20366040
P10-13
Argira Mapping Variants Series (Pipelines 10–13) — Test of hue→frequency functions
Software · Zenodo · May 24, 2026 · Log r=+0.878 · Mel r=+0.899 · Stochastic r=+0.935 · Quadratic r=+0.919 · N=30

Five pipelines testing log, mel, stochastic, and quadratic as hue→frequency functions, with correlations between r=+0.878 and r=+0.935, N=30.

10.5281/zenodo.20366726
SB
ARGIRA Synthetic Controlled Benchmark v2 — Independence of H, S, I (10 synthetic images)
Dataset · Zenodo · May 25, 2026 · Contrast 04 vs 05: same hue_std=0.5000, 6 vs 13 harmonics · Model A=f(H,S,I)

10 synthetic images isolating hue, saturation, and spatial irregularity; contrasts two images with the same hue_std but a different number of harmonics.

10.5281/zenodo.20385782
GE
ARGIRA Gradient Experiment v2 — Full acoustic response (21 images I00–I20)
Dataset · Zenodo · May 26, 2026 · Spatial irregularity gradient · per-level metrics + normalized curves figure

Acoustic characterization of 21 synthetic images (I00–I20) along a spatial irregularity gradient, with per-level metrics.

10.5281/zenodo.20387081
TN
Stereo Width: Malevich vs. Kandinsky — The inverted-complexity paradox
Technical note · Zenodo · May 2026 · W Malevich = 0.196 · W Kandinsky = 0.058 · Web Audio demo

Technical note on a stereo-width paradox: Malevich yields W=0.196 and Kandinsky W=0.058, a result inverted relative to what complexity would predict.

10.5281/zenodo.20353894
Infrastructure · Foundational software and corpora
v3.4
Argira Sonification Pipeline v3.4 — Empirical calibration, 3-orthogonal-channel architecture
Dataset · Zenodo · Jun 5, 2026 · N=187 · edge_density→mod_depth r=+0.900 (OLS) · roughness dropped for collinearity

Empirical calibration of a 3-orthogonal-channel architecture over N=187 artworks; edge_density→mod_depth r=+0.900 (OLS), with roughness dropped for collinearity.

10.5281/zenodo.20550714
B9
ARGIRA B_v9 — Spectro-Temporal Analysis of Sonified Visual Art (72 artworks)
Preprint · Zenodo · May 18, 2026 · 2,160 band-speed obs. · regime transition 4–8 kHz · ρ +0.973→−0.22

Spectro-temporal analysis of 72 sonified artworks, 2,160 band-speed observations, with a regime transition at 4–8 kHz where ρ goes from +0.973 to −0.22.

10.5281/zenodo.20266242
B9D
ARGIRA B_v9 Dataset — Structural Persistence Across Spectro-Temporal Transformations
Dataset · Zenodo · May 17, 2026 · 72 artworks · 6 bands × 5 speeds · ρ = −0.9524 (4–8 kHz vs hue_std)

Dataset of 72 artworks crossing 6 frequency bands and 5 speeds, with ρ = −0.9524 between the 4–8 kHz band and hue_std.

10.5281/zenodo.20258314
CC
Cycle Closure Experiment — Geometric Invariants for Sonified Visual Art (74 artworks)
Dataset · Zenodo · May 2026 · Inversion not possible (mean error 0.709) · cx → pan error < 0.05 · edge_density → effort r = 0.759

Test of geometric invariants over 74 artworks: full inversion of the process is not possible (mean error 0.709), though cx→pan has error <0.05 and edge_density→effort r=0.759.

10.5281/zenodo.20322821
D17
Argira Dataset v17 — 74 artworks · hue_std → Ds r = 0.9233
Dataset · Zenodo · May 2026 · 74 public-domain artworks · R² = 0.8524 · p < 0.001 (permutations N=1000)

74 public-domain artworks with hue_std→Ds r=0.9233, R²=0.8524, validated with a permutation test (N=1000, p<0.001).

10.5281/zenodo.20126917
P16
Argira Sonification Pipeline v16 — N=47, r=0.9289 (extended, not peer-reviewed)
Preprint · Zenodo · May 2026 · VTI → Ds r = 0.9289 · Sobel cross-validation r = 0.9570

Extended, not-peer-reviewed version, N=47: VTI→Ds r=0.9289, cross-validation with Sobel filter r=0.9570.

10.5281/zenodo.20123821
S15
Argira Station v15 — Standalone Analyzer (6 validated metrics)
Software · Zenodo · May 2026 · hue_std → Ds r = 0.9289 · irregularity → Sobel r = 0.9570 · N=47

Standalone analyzer with 6 validated metrics: hue_std→Ds r=0.9289, irregularity→Sobel r=0.9570, N=47.

10.5281/zenodo.20113356
v22
Argira Sonification Analyzer v22 — Object description with Claude Vision
Publication · Zenodo · May 13, 2026 · Complementary Claude Vision + Argira integration · proxy server (Node.js/Render) · fallback if the API is unavailable

Integrates Claude Vision as a complementary object-description channel, via a proxy server (Node.js/Render) with fallback if the API is unavailable.

10.5281/zenodo.20170517
v23
Argira Sonification Analyzer v23 — Spatial audio · centroid.x → pan · centroid.y → register
Software · Zenodo · May 16, 2026 · 4 perceptual dimensions · Web Audio API · active version of the sound museum

Adds spatial audio (centroid.x→pan, centroid.y→register) across 4 perceptual dimensions with the Web Audio API; this is the active version used by the sound museum.

10.5281/zenodo.20234082
v18–21
Argira Sonification Analyzer — Versions v18, v19, v20, v21
Software · Zenodo · May 2026 · Evolution: modular architecture → speed control → high-shelf filter → spatial position announcement

Four consecutive versions of the analyzer: v18 introduces modular architecture, v19 adds speed control, v20 adds tactile sonification with night mode, v21 adds spatial position announcement on a 3×3 grid.

v14
Argira Pipeline v14 — Dataset N=47, r=0.9289
Software · Zenodo · Apr 2026 · Permutation test N=1000 · 275 views · 140 downloads

Dataset of N=47 artworks with r=0.9289, validated with a permutation test (N=1000).

10.5281/zenodo.20097327
B9v2
ARGIRA B_v9 Dataset v2 — Threshold v* and client-side haptic module (43 artworks)
Dataset · Zenodo · May 2026 · v* = 7800 − 12500 · hue_std · haptics active in demo · r = 0.9370

Applies the threshold v* = 7800 − 12500·hue_std to 43 real artworks (r=0.9370) and activates the client-side haptic module in the demo.

10.5281/zenodo.20357570

v1 (10.5281/zenodo.20354217 ↗) was the technical note that formalized the v* = k − m·hue_std threshold model across the ecosystem's three layers (Python, JS, Android haptic extension), with no dataset of its own. v2 is the first version to apply the threshold to real data (43 artworks) and report the correlation.

Vis
Argira Vision v0.5 — Spoken semantic image description (Claude API)
Software · Zenodo · Apr 2026 · Complementary semantic channel · Web Speech API · no server

Adds a complementary semantic channel that describes images aloud using the Claude API and the Web Speech API, with no server required.

10.5281/zenodo.18989327
Web
ARGIRA Investiga — Interactive research page for the project (v3.0)
Software · Zenodo · Jun 4, 2026 · Experiments 1–15 · ARGIRA I–V findings · mapping taxonomy · WCAG AA

Interactive page bringing together experiments 1–15 and the findings of the ARGIRA I–V series, with a mapping taxonomy, conforming to WCAG AA.

10.5281/zenodo.20545664
v15
Argira Station v15 — Standalone analyzer, 6 validated metrics
Software · Zenodo · May 2026 · hue_std ↔ Ds r = 0.9289 · N=47 · canonical pipeline

Standalone analyzer with 6 validated metrics: hue_std↔Ds r=0.9289, N=47.

10.5281/zenodo.20113356
v16
Argira Sonification Pipeline v16 — Extended version (N=47, r=0.9289)
Preprint · Zenodo · May 2026 · Following the ICAD submission (v10) · Sobel r = 0.9570 · negative 2D FFT result documented

Extended version following the ICAD submission (v10), N=47, r=0.9289, Sobel r=0.9570; also documents a negative result with 2D FFT.

10.5281/zenodo.20123821
v17
Argira Dataset v17 — 74 artworks, r = 0.9233
Dataset · Zenodo · May 2026 · Independent replication with extended corpus · R² = 0.8524

Independent replication with the corpus extended to 74 artworks, r=0.9233, R²=0.8524.

10.5281/zenodo.20126917
08
Argira Scanline Pipeline v2 — Control without hue→frequency
Software · Zenodo · May 24, 2026 · r = −0.116 · Spatial architecture · Null result documented

Control pipeline with spatial architecture and no hue→frequency mapping; documents a null result (r=−0.116).

10.5281/zenodo.20366040
v18
Argira Sonification Analyzer v18 — Modular version, high-shelf filter
Software · Zenodo · May 2026 · Browser tool · Web Audio API · Modular architecture

Rewrites the analyzer with modular architecture and a high-shelf filter, as a browser tool built on the Web Audio API.

10.5281/zenodo.20133939
v19
Argira Sonification Analyzer v19 — Speed control (temporal magnifier)
Software · Zenodo · May 2026 · Range 0.5×–1.5× · Preserves correlation r > 0.92

Adds speed control (temporal magnifier) in the 0.5×–1.5× range, preserving correlation r>0.92.

10.5281/zenodo.20135805
v20
Argira Sonification Analyzer v20 — Tactile color sonification, night mode
Software · Zenodo · May 2026 · Pixel touch → pitch · Saturation → timbre · Brightness → volume

Tactile color sonification in night mode: pixel touch→pitch, saturation→timbre, brightness→volume.

10.5281/zenodo.20155401
v21
Argira Sonification Analyzer v21 — Spatial position announcement (3×3 grid)
Software · Zenodo · May 13, 2026 · Position before color · "upper left, blue"

Announces the position on a 3×3 grid first, followed by the color — for example, "upper left, blue."

10.5281/zenodo.20157205

From sonification
to communicated uncertainty

Starting in August 2026, the ARGIRA program extends to a distinct and complementary problem: how to communicate the uncertainty of an automatic alt-text estimate to a person who cannot verify it visually. Unlike commercial screen readers, which present their descriptions as facts, ARGIRA v1.4.x explicitly exposes the confidence of each estimate and the signals that support it — using deterministic image statistics, not neural networks.

This line includes three independent validation studies on the reliability of the technical signals used (stability of fractal_D, vulnerability of origin-discrimination features, and empirical RMA validation with 1,200 controlled runs), published with the same standard of methodological honesty as the rest of the program: what the signals can support is documented, as is what they cannot.

v1.4.7
ARGIRA v1.4.7 — Three-Layer Uncertainty Communication Prototype for Automatic Alt-Text New
Software · Zenodo · Aug 17, 2026 · Language selector · navigable card-based results · no changes to the classification algorithm

Three blocks of changes: (1) ES/EN language selector with cross-linking between versions; (2) result navigation via individual cards with explicit focus — no longer announced automatically on insertion; adaptations for JAWS (partially verified), TalkBack (inherited from earlier versions), and VoiceOver (adapted in code, not verified in a real environment); (3) control reordering, revised error text, removal of internal language from the header visible to the user. Introduces no changes to the classification heuristic or the uncertainty calculation. The two HTML artifacts (ES/EN) are self-contained: they work locally without installation or a network connection. The author reiterates that ARGIRA remains an experimental research artifact, not a finished product or validated accessibility technology.

10.5281/zenodo.21970403
Version history (v1.0 – v1.4.6)
v1.4.6 · Aug 16, 2026 Rewrite of the Layer 2/3 interface: clearer-language explanation, more prominent presentation of the instrument's estimate, visual indication of the runner-up classification, a separate notice when an estimate's reliability is reduced, reorganization of "More information" keeping each collapsible section tied to its control, and an A+/A- text-size control. Fixes a responsive accessibility bug: keyboard/screen-reader focus on a result card could end up hidden behind the fixed navigation bar; the fix adapts to the bar's actual height, even with a narrow viewport or the text-size control active. Verified with an automated test across 4 viewport widths (320/375/414/768px) plus the text-size control at 130%, waiting for scroll to stabilize before measuring (no arbitrary delays); 12/12 cases passed. TalkBack remains verified from earlier versions; VoiceOver was not evaluated due to lack of a test environment — untested scope, not a failed validation. No changes to the classification algorithm. Author's note: the DOI in the HTML footer mistakenly points to v1.4.4 instead of this version; this does not affect content or behavior. 10.5281/zenodo.21965506 ↗
v1.4.5 · Aug 16, 2026 Deliberately and stably fixes the public uncertainty-notice behavior: KNOWN_UNRELIABLE_MODE is now fixed to "brief automatic notice" (brief_auto_notice), shown whenever a result is classified as known_unreliable, identically in the English and Spanish HTML artifacts. The stated goal is for the public demo to communicate its current behavior directly and consistently, without turning demo users into subjects of an unannounced experiment. Reveals and documents the existence of an unpublished internal checkpoint, built on top of v1.4.4, that had linked that notice to SESSION_CONDITION, an A/B experimental condition (50/50 between "on-demand only" and "brief automatic notice") assigned per session on each page load — that A/B test was never published as such on Zenodo. The SESSION_CONDITION and assignSessionCondition() code remains present in the package but disconnected from public behavior, as infrastructure for a possible future experimental phase that would require its own methodology, data-collection plan, informed consent, and privacy design; this version adds no user feedback mechanism, experimental data collection, or mode selector. A brief, visible section is also added to both HTML artifacts explaining the currently active uncertainty-communication mode. It does not modify the image classification heuristic algorithm, the calibration table, the margin/runner-up calculation (0.18 threshold), level1PlusNotice(), the notice_mode_used field, or its CSV/JSON exports — everything remains the same as in the previous working checkpoint, as does the consistency fix between the on-screen text, the aria-label, and the copy/apply-to-alt-text actions and their exports. Accessibility verification: manual verification with Android TalkBack from earlier versions is retained; VoiceOver was not evaluated due to lack of a test environment — untested scope, not a failed validation. No external user study or evaluation with participants was carried out for this version. The package includes the English/Spanish artifacts, README, CHANGELOG, CITATION.cff, CC BY-NC-SA 4.0 license, and historical CSV exports from the experimental checkpoint alongside new exports confirming notice_mode_used consistently as brief_auto_notice under the fixed v1.4.5 behavior. It remains an experimental research artifact, not presented as a finished product or as validated accessibility technology. 10.5281/zenodo.21959360 ↗
v1.4.4 · Aug 13, 2026 Accessibility and localization refinement release. Reviews and refines English-language communication across the main interface: result information, accessibility labels, dynamic messages, metadata panels, and explanatory content, keeping the Spanish variant aligned with the English one — both remain independent HTML artifacts, with no in-app language selector. Preserves the three-layer uncertainty-communication approach and the already-established distinction between the capabilities of this specific demo (statistical pixel heuristics, with no computer-vision model with semantic recognition) and the broader ARGIRA project; the generated text continues to be described as an automatic type description, not a semantic description of the image. Also preserves the established navigation in accessibility-only mode and the essential result flow: image type, confidence, generated alt-text, copy, and previous/next navigation. Accessibility verification: manual verification with Android TalkBack from earlier versions is retained; VoiceOver was not evaluated due to lack of a test environment — untested scope, not a failed validation; no external user study was carried out. Introduces no new image classification algorithm, calibration model, or experimental condition — the heuristic classification, calibration logic and thresholds, embedded data, identifiers, and experimental session mechanism are preserved unchanged from v1.4.3. It remains an experimental research artifact, not presented as a finished product or as validated accessibility technology. 10.5281/zenodo.21924218 ↗
v1.4.3 · Aug 13, 2026 Accessibility and localization release. Refines the introductory explanation of what the prototype does and, importantly, explicitly distinguishes the capabilities of this specific demo — statistical pixel heuristics and predefined templates, with no computer-vision model with semantic recognition; it can estimate an image's probable type and expose its statistical signals and uncertainty, but it does not identify objects, people, or scenes — from the broader ARGIRA project, which does include other tools and experimental approaches for image content; this prevents the limitation of this specific prototype from being mistakenly attributed to ARGIRA as a whole. It also clarifies that the generated text is an automatic type description, not a semantic description of the image. The browser CORS limitation notice changes from an always-expanded block to a native disclosure (details/summary pattern), reducing visual and navigational noise when opening the app without hiding the information from anyone who needs it; it remains excluded from the reduced path in accessibility-only mode. Reviews interface text consistency, accessible labels, dynamic messages, metadata panels, and explanatory content in both language variants, which remain independent HTML artifacts. Accessibility verification: manual verification with Android TalkBack from earlier versions is retained; VoiceOver was not evaluated due to lack of a test environment — untested scope, not a failed validation; no external user study was carried out. Introduces no new classification algorithm, calibration model, or experimental condition — heuristic classification, calibration logic, thresholds, embedded data, identifiers, and the experimental session mechanism are preserved unchanged from v1.4.2. It remains an experimental research artifact, not presented as a finished product or as validated accessibility technology. 10.5281/zenodo.21916695 ↗
v1.4.2 · Aug 8, 2026 Accessibility-focused patch. Refines accessibility-only mode to reduce unnecessary screen-reader navigation: sample-image notices and notes are now hidden from the accessibility path, along with the "More details"/"More information" controls and their associated content, leaving the essential flow available (image type, confidence, generated alt-text, copy, previous/next navigation). Also fixes an actual bug in the previous implementation: the CSS class that hid the collapsible panel did not reach the button that opened it or the separate "More information" button, so both remained reachable by screen reader even though their panels were hidden. Manually verified with Android TalkBack in both the English and Spanish variants. The six bulk export actions (copy all as text/CSV/JSON and download TXT/CSV/JSON) are now grouped under a single expandable native "Export options" control, available in both normal mode and accessibility-only mode; when collapsed, the six actions no longer add separate stops in screen-reader navigation. The underlying export functions do not change. Also corrects the number of sample images included (from four to three) and shortens the accessibility-only mode description to reduce redundant screen-reader output. Verification carried over from v1.4.1: it remains unresolved that the "More information" section should open on activation rather than merely on receiving focus; VoiceOver was not evaluated due to lack of a test environment — untested scope, not a failed validation; no external user study was carried out. Introduces no new classification algorithm, calibration model, or experimental condition; the experimental session mechanism, identifiers, classification logic, calibration logic, thresholds, and data flow are preserved. Structural parity maintained between the English and Spanish artifacts. 10.5281/zenodo.21852709 ↗
v1.4.1 · Aug 8, 2026 Visual and touch-target accessibility patch, CSS only. Increases base interface text from 14px to 16px, generated alt-text to ~16px, primary controls to ~14px, and secondary controls to ~13px with touch-target height increased from ~28px to ~38px; privacy/status text and result metadata move to ~13px; the image-type badge keeps its compact scale as a brief status label, not primary content. The change was applied identically to the English and Spanish artifacts, which end up with byte-for-byte identical stylesheets. No changes to HTML structure, element identifiers, JavaScript, classification logic, calibration logic, experimental session mechanism (SESSION_ID, SESSION_CONDITION), thresholds, or data flow. Verification carried over from v1.4.0: manual verification with Android TalkBack confirmed 8 of 9 checklist points in the forced experimental variants used in Fase 3C; it remains unresolved that the "More information" section should open on activation rather than merely on receiving navigation focus. VoiceOver was not evaluated due to lack of a test environment — untested scope, not a failed validation; no external user study was carried out. Both artifacts derive from the frozen v1.3.0-alpha prototype and preserve the same experimental session mechanism, identifiers, classification logic, calibration logic, thresholds, and data flow, with structural parity maintained between variants. 10.5281/zenodo.21850529 ↗
v1.4.0 · Aug 8, 2026 Incremental release extending the v1.3.0-alpha prototype on two fronts: accessibility verification and Spanish localization. Includes the developer's first manual accessibility verification with Android TalkBack on the forced experimental variants used in Fase 3C, following the project's predefined screen-reader checklist: 8 of 9 points confirmed; it remains explicitly unresolved that the "More information" section should open only on activation, not simply on receiving navigation focus. Apple VoiceOver was not evaluated due to lack of a test environment — untested scope, not a failed validation. Adds a Spanish variant of the prototype (argira_demo_es.html) alongside the original English artifact (argira_demo_v2.html), as two independent, separately runnable HTML artifacts; language selection in this version is done by choosing the corresponding artifact — there is no in-app language selector, no switcher, no automatic detection based on browser language, and no cross-linking between variants. Both variants derive from the same frozen v1.3.0-alpha prototype and preserve the same experimental session mechanism, identifiers, classification logic, calibration logic, thresholds, and data flow; localization covers the already-defined interface and generated-content perimeter, including the accessibility description function and the audited generated-text templates — where English uses language-specific article logic, Spanish uses native grammatical construction while preserving the same semantic outcome. Structural verification confirmed parity between both artifacts, including JavaScript function structure, element identifiers, SESSION_ID, and SESSION_CONDITION; the differences are restricted to the authorized localization content and Spanish-specific grammatical construction. Introduces no new classification algorithm, calibration model, experimental condition, or language-selection interface. 10.5281/zenodo.21849334 ↗
v1.3.0-alpha · Aug 7, 2026 Extends the v1.2 infrastructure toward a future exploratory accessibility evaluation. The analytical pipeline remains unchanged from v1.2: for the same input image, classification, calibration assessment, confidence estimation, and the generated alt-text do not differ between the two versions. Introduces session- and event-level instrumentation: session identifiers attached to results and lightweight logging of interactions with the uncertainty-notice controls. This infrastructure is intended to support a future accessibility evaluation with screen-reader users; no user study or comparative analysis was carried out in this version. It is presented as a research prototype for documentation and experimentation, not as a production-validated image classification system; the Layer 2 and Layer 3 validation carried out for v1.1 was not repeated here because the validated analytical pipeline did not change. The Result Manager (Layer 3.5), session instrumentation, and event instrumentation are considered infrastructure components, not analytical components, and have not themselves been the subject of a formal validation study. 10.5281/zenodo.21843103 ↗
v1.2 · Aug 7, 2026 Extends the demonstration and result-management infrastructure while preserving the validated analytical pipeline from v1.1; the Layer 1–2.5 heuristics do not change. For the same input image, neither classification, calibration assessment, confidence estimation, nor the generated alt-text differ between v1.1 and v1.2. Introduces Layer 3.5 (Result Manager): in-memory result persistence, multi-format export (TXT, CSV, and JSON) through a single serialization interface, copy to clipboard, and file download for each supported format, plus accessibility improvements including a skip link and explicit ARIA relationships between collapsible controls and their associated panels. Published as a research prototype for documentation and experimentation, not as a production-validated image classification system; the Layer 2 and Layer 3 validation carried out for v1.1 was not repeated in this version because the validated analytical pipeline did not change. The new Result Manager (Layer 3.5) is considered an infrastructure component, not part of the analytical pipeline, and has not itself been the subject of a formal validation study. 10.5281/zenodo.21839157 ↗
v1.1 · Aug 7, 2026 First implementation of the three-layer communication architecture, in both Python and JavaScript. Both implementations were verified to produce the same Layer 2 communication result across all supported combinations of calibration status/predicted type, and the browser prototype was manually validated. The package includes complete validation material (architecture specification, calibration table, knowledge model, wording reviews, and a validation closure document), the Python reference implementations and the browser prototype, an earlier reference implementation with the sample images used in validation, SHA-256 checksums of the entire package for integrity verification, and a section of open work on heuristic parity between the Python and JavaScript Layer 1 implementations — explicitly unresolved and not part of the validation completed for this version. Published as a research prototype for documentation and experimentation, not as a production-validated image classification system. 10.5281/zenodo.21834198 ↗
v1.0-Experimental · Aug 7, 2026 First frozen experimental release of ARGIRA's image-type classification instrument. The repository contains the original HTML heuristic, a faithful Python implementation used exclusively for characterization, characterization results on a 162-image corpus, methodological documentation, and reproducibility materials. The heuristic classifies images into four predefined categories (painting, photograph, AI-generated, illustration) using deterministic image statistics (irregularity, entropy, edge density, hue entropy, and saturation entropy); it uses no machine-learning models or trained classifiers. The characterization corpus contains 108 paintings, 15 photographs, 19 synthetic geometric patterns, and 20 real AI-generated images made with Stable Diffusion obtained from Lexica. Main characterization results: 38.9% overall accuracy; 0% detection of the real AI-generated images included in the corpus; no high-confidence errors (confidence ≥ 0.75). This release documents the instrument exactly as it was evaluated — it is not presented as a state-of-the-art AI-image detector or as a benchmark. The original corpus images are not redistributed; the repository includes only hashes and metadata to support reproducibility while respecting image licenses. Future improvements will be published as new versions, preserving this frozen baseline for reproducible research. 10.5281/zenodo.21830302 ↗
Val
Stability of ARGIRA's Fractal-D Feature and RMA-0 Audit — Validation and Reproducibility
Dataset · Zenodo · Aug 15, 2026 · 52-image corpus · 7 experimental phases

Study complementary to ARGIRA v1.4.4 investigating whether the observed behavior of the fractal_D feature remains stable under changes in scale, interpolation method, intermediate resolutions, and controlled transformations, together with an audit of the associated residual (RMA-0) — distinct from the independent RMA calibration repository (1,200 runs, DOI 10.5281/zenodo.21935413). Experimental corpus of 52 images (18 AI-generated, 14 photographs, 20 paintings). Chains seven documented phases: fractal_D stability across the full corpus, analysis of individual D curves and successive differences, Lanczos interpolation control at different scales, bilinear vs. Lanczos comparison, intermediate-resolution analysis, residual transformation analysis, and the RMA-0 audit with control for image format/origin effects. The central question is whether the fractal_D signal is a stable property of the analyzed images or whether its measured value can be substantially affected by scale, interpolation, resolution, or controlled transformations — thus examining the signal itself and the conditions under which it changes, not an isolated numerical value as independent evidence of origin. The study does not claim that fractal_D alone determines the origin of an individual image, nor does it constitute a complete system for detecting AI-generated images; the results should be interpreted within the documented corpus, implementations, transformations, and experimental conditions. A 100-file reproducibility package (99 covered by an MD5 integrity manifest) organized into scripts, CSV/JSON numerical data, figures, reports, the 52-image corpus, the materials from the format/origin control phase, the RMA-0 audit materials, and the original source of box_counting_dimension included as a methodological reference to verify the fidelity of the analyzed implementation. The results are intended to inform ARGIRA's future development, including the planned evolution toward v1.4.5, but this deposit does not itself constitute the implementation of that version — publishing the research and its eventual incorporation into the software are separate steps. Complementary to, and not overlapping with, the independent study on the vulnerability and controllability of origin-discrimination signals (DOI 10.5281/zenodo.21945633), which focuses on luminance and saturation signals rather than fractal_D.

10.5281/zenodo.21946989
Val
Vulnerability and Controllability of Origin-Discrimination Features (AI / Painting / Photograph)
Dataset · Zenodo · Aug 15, 2026 · 9 experimental phases · documented negative result for fractal_D

Study complementary to ARGIRA v1.4.4 investigating whether ARGIRA-specific signals can reliably distinguish between AI-generated images, real paintings, and real photographs. Main result: the signals studied do not constitute reliable evidence of origin under the experimental conditions analyzed. fractal_D does not discriminate significantly between AI-generated images and paintings when using the real production implementation on a balanced sample; luminance_contrast + sat_mean does show statistical separation between groups, but it is vulnerable to routine brightness, contrast, and saturation edits; the production score z_combined can be shifted through controlled photo edits without the image's structural features shifting in the same way; and stress tests show within-class score variability comparable to the separation observed between classes. Chains nine studies or experimental phases with verifiable artifacts: exploratory test of fractal_D and candidate features culminating in a Ronda 5 with 20 AI-generated images versus 20 paintings using a faithful port of the real ARGIRA v3.5 production implementation; separation with luminance_contrast + sat_mean; validation with development and validation splits; cross-validation with Gemini-generated images; experimental control for lighting, brightness, contrast, and saturation; luminance_mean audit; vulnerability control with real paintings; systematic stress test with 13 perturbation families on three base images; and an extended stress test with nine base images including AI-generated, paintings, and photographs. The natural-pairing and edit-based pairing phases (10C/10A) are not included as reproducible results because no corresponding verifiable artifacts were located. The study does not establish the origin of any individual image with certainty, does not constitute a complete AI-image detection system, does not evaluate every existing generator or generation pipeline, does not evaluate machine-learning architectures as alternative approaches, nor does it establish that the observed behavior necessarily generalizes to all possible datasets, generators, or image conditions — these limitations do not invalidate the findings, but they do delimit their scope. It is an independent scientific repository, not part of the ARGIRA v1.4.4 release itself; its results are intended to inform the evolution toward v1.4.5, but the publication of this study and its implementation in the HTML are separate decisions — the scientific evidence is first established and published in a traceable way, and only afterward is it decided which specific conclusions are incorporated into that future version's interface and logic, which should cite this repository by its DOI as part of the scientific traceability of the resulting changes.

10.5281/zenodo.21945633
Val
Empirical Validation of RMA Under Synthetic Control Conditions
Dataset · Zenodo · Aug 14, 2026 · 1,200 runs · 0% false positives in pure-noise control

Methodological validation/calibration study carried out as part of the development of version 1.4.5, continuing the RMA (Residual Mapping Analysis) analyses used in the preceding v1.4.4. Documents an empirical validation of RMA under synthetic control conditions with known ground truth, assessing its ability to discriminate between correctly specified or signal-free data and deliberately misspecified models. Five generative conditions were tested: an external pure-noise control, a correctly specified linear condition, and three deliberately misspecified conditions (nonlinear, composite, and hierarchical); each condition was evaluated at four sample sizes (n = 51, 100, 300, and 800), with 30 independent random seeds and two targets, resulting in 1,200 final runs. Under the tested conditions, correctly specified and signal-free cases produced negative residual R² values converging toward zero as sample size increased, while misspecified conditions produced positive residual R² values with increasing separation as the sample grew; the pure-noise control produced 0% observed false positives across 240 runs under the framework's default threshold (0.05). The study provides empirical support for RMA's discriminative behavior under the synthetic conditions tested, but it does not constitute a general validation of RMA for all possible configurations, nor does it validate downstream findings that use RMA in D or in ARGIRA. The repository includes complete raw results, aggregated results, the calibration script, the block-execution wrapper, the frozen experimental protocol, README documentation, and the main visualization; the final dataset contains 1,200 rows with no duplicate experimental keys and a framework integrity hash consistent across all records. Version v1.1.0 (current, linked here) adds rma_framework.py, the reference implementation of the RMA audit pipeline used to produce these results, mistakenly omitted from the earlier v1.0.0 — a code-completeness update: the underlying data, results, and figures do not change; MANIFEST.md5 and CITATION.cff were regenerated to reflect the new file and version.

10.5281/zenodo.21935413
v1.4.4
ARGIRA v1.4.4 — Three-Layer Uncertainty Communication Prototype for Automatic Alt-Text
Software · Zenodo · Aug 13, 2026 · Accessibility and localization refinement · see full history above (v1.0 – v1.4.6)
10.5281/zenodo.21924218
Vis
Argira Vision v0.5 — Spoken Image Description for Blind People (Claude API)
Software · Zenodo · Local history · Full ARIA · Haiku/Sonnet selection · voice adjustment
10.5281/zenodo.19759118

Concept DOI for this line (always resolves to the latest version): 10.5281/zenodo.21830301 · Full tool-ecosystem annex: view annex ↗


Jose Ranero García

Principal Investigator ARGIRA Project 2026

ARGIRA is an independent research project on sonification and audiovisual perception, published openly under a CC BY-NC 4.0 license. The project documents both positive and negative findings, in the conviction that mapping the limits of knowledge is as valuable as extending them.


Complete materials —scripts, datasets, figures, and preprints— are available on Zenodo. Negative results are results.