Research Data and Reproducibility Policy

Research Data and Reproducibility Policy

JSOMER promotes transparent, verifiable, reproducible, and ethically responsible research while recognizing legitimate limits on public data sharing.

Purpose

This policy establishes requirements and recommendations for research data, Data Availability Statements, repositories, metadata, citation, code, materials, protocols, computational reproducibility, qualitative and mixed-methods research, social media and platform data, health and clinical information, controlled access, third-party data, documentation, retention, editorial verification, corrections, and data-integrity concerns. Requirements may vary by article type, discipline, design, consent, ethics, law, sensitivity, intellectual property, contracts, and platform conditions.

Data and Reproducibility at a Glance

Data Availability Statement requiredResearch and evidence-synthesis manuscripts must state whether supporting data exist, where they are located, and how they may be accessed.
Share when appropriateData, code, materials, and protocols should be shared when ethically, legally, contractually, and technically possible.
Restrictions must be explainedPrivacy, consent, platform rules, community rights, security, third-party ownership, and law may justify controlled or restricted access.
No automatic open-data requirementA manuscript is not unsuitable merely because public sharing is inappropriate, provided the restriction is legitimate and transparent.
Supporting evidence may be requestedEditors may request data, code, records, or confidential expert access when necessary for evaluation or an integrity inquiry.
Authors retain responsibilityRepository deposit and peer review do not transfer responsibility for data accuracy, documentation, ethics, or integrity to the journal.
How to read this policy: The principal data-sharing and reproducibility requirements remain visible above. Select any numbered section below to expand the detailed standards and model statements. All provisions remain in the page source when a section is closed.

Data Availability, Documentation, Access, and Verification

01General Principles and Data-Sharing Model View details +

JSOMER supports transparent, verifiable, reproducible, and ethically responsible research. Research claims should be supported by appropriately documented evidence, and methods and analytical decisions should be described clearly enough for scholarly assessment.

The journal’s model is: share data, code, and research materials when ethically, legally, contractually, and technically possible. When sharing is restricted, explain the restriction clearly and provide the most transparent and workable form of access that remains appropriate.

  • Data Availability Statements must accurately describe whether supporting data exist, where they are located, and how they may be accessed.
  • Privacy, consent, ethics approval, community rights, intellectual property, contracts, platform conditions, security, and applicable law take priority over unrestricted sharing.
  • The inability to share data publicly does not automatically make a manuscript unsuitable for publication.
  • Editors and reviewers may request reasonable supporting evidence when needed for evaluation or an integrity inquiry.
  • Data and materials must not be fabricated, falsified, deceptively manipulated, improperly withheld, or misrepresented.
02FAIR Principles and Key Definitions View details +

Where appropriate, authors are encouraged to make research objects findable, accessible, interoperable, and reusable. FAIR does not mean that every dataset must be openly available. Restricted or controlled access may be appropriate, and even nonpublic data may have public metadata describing their existence and access conditions.

Research data include information, observations, measurements, records, media, text, numerical values, outputs, and other evidence supporting a study. Raw data are the original collected or generated records; processed data have undergone documented cleaning, coding, transformation, derivation, or aggregation; source data are the records from which published findings were derived; and metadata describe origin, structure, context, ownership, and access conditions.

Research materials may include instruments, interview guides, stimuli, protocols, codebooks, search strategies, extraction forms, data dictionaries, prompts, model cards, and software environments. Analytical code includes scripts, notebooks, workflows, algorithms, queries, and instructions used to prepare, analyze, visualize, or model data. Computational reproducibility concerns obtaining substantively consistent results using the same or equivalent data and procedures; replication concerns a new study using new data, participants, settings, or conditions.

03Data Availability Statement Requirements View details +

Every manuscript reporting original research, secondary data analysis, systematic evidence synthesis, meta-analysis, computational analysis, or methodological validation must include a Data Availability Statement. A statement is required even when no data can be shared.

  • whether data were created or analyzed;
  • whether data are available and under what conditions;
  • repository name, dataset title, persistent URL, DOI, accession number, or version where applicable;
  • whether access is open, embargoed, restricted, controlled, or available by request;
  • who may apply, how to request access, and what approvals or agreements are required;
  • availability of analytical code, documentation, protocols, or research materials; and
  • a meaningful reason when public sharing is not possible.

The statement must not be vague, misleading, inconsistent with the actual materials, or dependent solely on an unstable personal webpage, temporary transfer service, or private cloud link.

04Standard Data Availability Statements View details +
Open repository

Data Availability: The data supporting the findings of this study are openly available in [repository name] at [DOI or persistent URL] under the [license name] license.

Data and code available

Data and Code Availability: The data, analytical code, and documentation supporting the findings of this study are openly available in [repository name] at [DOI or persistent URL].

Included with the article

Data Availability: All data supporting the findings of this study are included in the article and its supplementary materials.

Controlled access

Data Availability: The data are not publicly available because they contain sensitive or potentially identifiable information. Qualified researchers may request controlled access from [authorized body]. Access is subject to [ethics approval, institutional authorization, data-use agreement, or other conditions].

Third-party data

Data Availability: The data were obtained from [provider] under a restricted-use agreement and cannot be redistributed by the authors. Researchers may request access directly from [provider and URL], subject to the provider’s requirements.

Social media data

Data Availability: Raw social media data are not publicly shared because redistribution could conflict with platform conditions, user privacy expectations, and restrictions concerning identifiable content. The manuscript and supplementary materials describe the collection, processing, and analytical procedures.

No new data

Data Availability: No new data were created or analyzed in this study. Data sharing is not applicable.

Qualitative data restricted

Data Availability: The qualitative data are not publicly available because full transcripts could permit participant identification and public sharing was not authorized by the consent process. Deidentified excerpts supporting the analysis are presented in the article.

05Requests, Controlled Access, and Embargoes View details +

The phrase “available upon reasonable request” must identify the data holder, request address, applicant eligibility, required information, permitted analyses, approval and agreement requirements, review process, grounds for refusal, possible costs, and expected period of availability. The corresponding author should not be the sole access point unless the author has durable authority and capacity to manage access.

Controlled access may use a data-access committee, institutional office, secure data environment, confidentiality agreement, ethics approval, or limited extract. Legitimate restrictions should be described as transparently as possible, including the authority imposing the restriction, whether aggregate data, metadata, code, synthetic data, or demonstration materials can be supplied, and whether the restriction may expire.

A temporary embargo may be accepted for repository processing, intellectual-property review, funder or contractual conditions, participant protection, or another legitimate reason. The statement should provide the reason, expected release date, future deposit location, post-embargo conditions, and any interim access procedure. Embargoes must not be used to prevent scrutiny indefinitely.

06Repositories, Persistent Identifiers, and Versioning View details +

When sharing is appropriate, authors should use a recognized institutional, disciplinary, or general-purpose repository providing suitable preservation, metadata, access controls, persistent identifiers, versioning, licensing, discoverability, governance, and long-term accessibility. Repositories issuing DOIs or established accession numbers are preferred.

Data, code, and materials may be deposited before or during review using private, anonymous, or controlled reviewer links when necessary. Double-blind review links should not reveal author identity unnecessarily. Authors must not replace, delete, or materially alter materials under review without informing the editor.

The exact version supporting the article must remain identifiable. Updates should preserve the original version, document material changes, and provide new version information. A correction should be considered when a post-publication change affects reported findings.

07Data Citation, Documentation, and File Formats View details +

Datasets and other substantial research objects should be cited appropriately. Citations should identify the creator, year, title, version, repository or publisher, persistent identifier, and access date where relevant. Data citation does not replace lawful and ethical access.

Shared data should include sufficient documentation for a qualified researcher to understand collection methods, sampling, eligibility, variables, units, coding, missing values, transformations, derived variables, cleaning, exclusions, quality checks, file relationships, software requirements, restrictions, and known limitations. A README and, for complex data, a data dictionary, codebook, schema, or metadata record are strongly encouraged.

Authors should use open, nonproprietary, or widely supported file formats when appropriate. Proprietary formats may be included when necessary, but an accessible alternative should be supplied where reasonably possible. Files should not depend on obsolete or inaccessible software without explanation.

08Analytical Code, Software, and Computational Environments View details +

Authors should share code supporting reported findings when ethically, legally, contractually, and technically possible. Documentation should identify the programming language, software and versions, dependencies, installation instructions, execution order, inputs, outputs, random seeds, parameters, preprocessing, analysis and visualization scripts, limitations, and license.

For complex computational work, authors are encouraged to document the operating system, hardware, memory, package-lock or environment files, containers, notebooks, workflows, cloud services, model checkpoints, random seeds, and expected runtime. Proprietary software use should be described with enough detail for informed evaluation.

Authors must not claim that code is available when deposited files are partial, nonfunctional, unrelated, or insufficient to support the stated analyses. When exact reproduction is impossible because of changing APIs, proprietary infrastructure, nondeterminism, unavailable platforms, or hardware dependence, the limitation should be explained explicitly.

09Artificial Intelligence, Machine Learning, and Generative AI View details +

AI and machine-learning studies should report data provenance, training, validation, and test-set definitions, preprocessing, architecture, model and software versions, parameters, hyperparameters, optimization, seeds, metrics, baselines, class imbalance, bias or fairness assessment, error analysis, model selection, human review, external tools or APIs, computational resources, release restrictions, and available code, weights, or checkpoints where appropriate.

When using commercial or externally hosted models, authors should identify the provider, model name and version where available, access date, material settings, prompts or instructions, known reproducibility limits, and whether the provider may change the system over time. Exact reproducibility must not be claimed when the model, API, training data, or environment cannot be preserved or accessed.

When generative AI is studied or used materially, authors should preserve appropriate prompts, available system instructions, conversation sequence, model identity, dates, sampling settings, output-selection procedures, number of generations, human editing, exclusions, safety filters, API settings, and evaluation procedures. Confidential, personal, copyrighted, illegal, or unsafe content must not be exposed through public sharing.

10Research Materials, Instruments, and Protocols View details +

Authors should share questionnaires, scales, interview and focus-group guides, stimuli, intervention manuals, educational materials, protocols, codebooks, coding frameworks, search strategies, extraction forms, preregistrations, analysis plans, and other methodological resources when lawful and appropriate.

For measurement instruments, authors should report the name, source, version, language, translation or adaptation process, scoring, relevant psychometric information, permissions, licenses, and study-specific modifications. Copyrighted instruments must not be reproduced in full without authorization.

Authors should state whether a protocol, preregistration, registered report, clinical-trial record, or analysis plan exists and cite or link to it where available. Prospective registration must precede the relevant observation or analysis. Retrospective registration must not be represented as preregistration, and material deviations must be identified and justified.

11Qualitative and Mixed-Methods Research View details +

Complete qualitative datasets are not always suitable for open sharing because anonymization may be difficult or may destroy context. Authors should consider consent, identifiability, sensitivity, group disclosure, third-party information, cultural expectations, legal restrictions, researcher-participant relationships, and risks arising from quotation or reuse.

When full data cannot be shared, authors should provide as much transparency as ethically possible through deidentified excerpts, coding frameworks, codebooks, theme definitions, analytic memos, reflexivity statements, methodological descriptions, controlled access, or other supporting documentation. The absence of public qualitative data does not itself prevent publication.

Mixed-methods studies should explain the relationship, sequence, and integration of quantitative and qualitative components, transformation procedures, joint displays, inconsistent findings, access conditions for each data type, and reasons for different sharing arrangements. Public quantitative data do not justify release of identifiable qualitative material from the same participants.

12Social Media and Internet-Derived Data View details +

Raw social media or internet-derived data should not be shared without careful consideration of consent, contextual privacy, usernames, searchable quotations, geolocation, sensitive disclosures, private messages, deleted content, group membership, vulnerability, stigmatization, platform and API terms, copyright, data-protection law, linkage risk, and downstream misuse.

Removing usernames alone may not anonymize content when users can be identified through quotations, timestamps, URLs, networks, images, profiles, locations, or linkage with public information. Usernames, handles, profile URLs, photographs, private messages, identifiable screenshots, closed-group content, sensitive exact quotations, personal networks, and location histories should not ordinarily be deposited publicly.

Where exact content is required for reproducibility, authors should consider permission, controlled access, paraphrasing, transformed identifiers, secure enclaves, platform-compliant methods, lawful content identifiers, synthetic examples, aggregate data, or detailed analytical documentation without raw redistribution. Technical access through a platform or API does not create an unrestricted right to redistribute content.

13Health, Clinical, Genetic, and Vulnerable-Population Data View details +

Health and clinical data must not be shared openly when doing so would violate consent, patient privacy, ethics approval, confidentiality, law, institutional policy, or data-use agreements. Deidentification claims must account for rare diagnoses, dates, locations, clinical histories, genetic information, free text, images, demographic combinations, linked data, and other indirect identifiers.

Clinical-trial reports must include a data-sharing statement describing whether deidentified participant data, dictionaries, protocols, analysis plans, and related documents will be shared; what will be shared; timing and duration; applicant eligibility; permitted analyses; review procedures; agreements; and access location.

Genetic, genomic, facial, voice, gait, behavioral, location, and other biometric data require heightened protection because implications may extend to relatives or communities. Data concerning children and vulnerable populations should be openly shared only when consent or authorization, ethics approval, protection, low reidentification risk, and assessment of downstream harm support the release.

14Community-Governed, Environmental, Third-Party, and Proprietary Data View details +

Research involving Indigenous peoples, identifiable communities, culturally sensitive knowledge, or collective interests may require community permissions, data sovereignty, collective privacy, culturally appropriate access, benefit sharing, limits on secondary use, attribution, and community decisions about repositories and licenses. FAIR practices must not override community rights.

Location-sensitive data concerning endangered species, fragile habitats, archaeological resources, or vulnerable communities may require spatial aggregation, coordinate masking, restricted access, generalized locations, delayed release, or removal of sensitive fields.

Third-party and proprietary data must be used in accordance with access conditions, licenses, consent, ethics, confidentiality, laws, repository rules, contractual restrictions, citation requirements, and redistribution prohibitions. Authors must identify the source and lawful access route and must not represent restricted data as available. Contracts that prevent independent analysis, interpretation, or publication may make a manuscript unsuitable for JSOMER.

15Synthetic, Simulated, Review, and Methodological Data View details +

Synthetic data must be identified clearly and accompanied by information about generation methods, contribution of real data, models, parameters, validation, similarity to real records, disclosure risk, limitations, and available code. Synthetic data are not automatically anonymous and must not be represented as observed empirical data.

Simulation studies should report the data-generating process, parameters, distributions, sample sizes, replications, assumptions, seeds, software, scenarios, performance metrics, code where possible, and limitations. Simulated findings must be distinguished from observations of real persons, populations, platforms, institutions, or environments.

Systematic reviews, scoping reviews, and meta-analyses should share complete search strategies, dates, sources, screening and extraction materials, study lists, exclusion reasons where required, extracted data, risk-of-bias judgments, effect-size calculations, code, sensitivity analyses, flow diagrams, and protocol or registration information where permitted. Copyrighted full texts must not be redistributed without authorization.

Methodological studies should provide sufficient data, code, simulations, examples, instruments, or documentation for meaningful evaluation. Where proprietary data or software are essential, accessible demonstrations, simulated examples, or pseudocode should be provided when possible.

16Reporting Transparency, Null Findings, and Data Reuse View details +

Authors must report sampling, recruitment, eligibility, exclusions, sample-size justification, measures, reliability and validity, cleaning, missing-data handling, transformations, outliers, model specifications, stopping rules, multiple comparisons, qualitative coding, reflexivity, preprocessing, bot detection, network construction, sensitivity analyses, model-selection decisions, and limitations as relevant to the design.

Findings must be reported accurately regardless of statistical significance, direction, novelty, or consistency with expectations. Authors must not suppress null findings, omit prespecified outcomes without explanation, switch primary outcomes silently, present exploratory work as confirmatory, exaggerate weak evidence, or treat inconclusive results as definitive.

When a dataset supports multiple publications, each manuscript must identify the dataset, cite related outputs, explain overlap, distinguish the new contribution, avoid presenting reused data as newly collected, avoid duplicate results, and comply with consent, ethics, and data-use terms.

17Data Retention, Provenance, Quality, and Secure Destruction View details +

Authors should retain data and essential supporting records for periods required by law, ethics approval, institutional and funder policy, professional standards, contracts, and the nature of the research. Where no longer period applies, retention should normally continue for at least five years after publication, with longer periods where appropriate for trials, longitudinal or regulated studies, reusable datasets, patents, investigations, or long-term consequences.

Records may include raw and processed data, code, dictionaries, consent and ethics documentation, protocols, registrations, plans, field or laboratory records, interview documentation, provenance, cleaning logs, exclusions, model outputs, version histories, access correspondence, data-use agreements, and support for tables and figures.

Authors should document data origin, collection date, access method, version, filtering, transformations, linkage, recoding, aggregation, derived variables, software, responsible personnel, and output location. Quality processes may include calibration, validation, duplicate and range checks, coding verification, interrater agreement, audit trails, source checks, missing-data assessment, sensitivity analysis, and independent verification.

When data must be destroyed under an approved protocol or binding requirement, destruction should be secure and documented, including what was destroyed, when, under what authority, which copies were covered, whether backups or derived data remain, and whether verification is affected.

18Editorial and Reviewer Access to Supporting Materials View details +

Editors may request data, code, materials, original media, documentation, or additional analyses to assess reliability, resolve inconsistencies, verify figures or tables, examine methodology, investigate selective reporting or possible fabrication, verify ethics statements, or respond to post-publication concerns. Reviewers must route requests through the handling editor.

When open sharing is not possible, confidential or controlled access may be provided to the editor, a statistical or methodological reviewer, a research-integrity adviser, or another qualified specialist through secure data rooms, confidentiality agreements, institutional approval, limited extracts, remote analysis, or another safeguard. JSOMER will not normally take custody of sensitive raw data unless a specific secure procedure has been established.

Normal peer review does not guarantee full audit of every dataset or execution of all code, and repository deposit does not establish validity. Authors remain responsible for accuracy, completeness, documentation, ethical handling, and integrity of data and code.

19Data Integrity, Corrections, and Post-Publication Concerns View details +

Fabrication includes inventing data, people, observations, records, interviews, platform content, images, repository records, or findings. Falsification includes unauthorized alteration, deceptive exclusion, changed dates, manipulated images, modified code, concealed contrary findings, altered outputs, or misrepresentation of collection methods. Legitimate cleaning, correction, transformation, exclusion, or technical adjustment is acceptable when justified, documented, consistently applied, and reported transparently.

Authors must notify JSOMER promptly when errors are found in data, code, analyses, figures, tables, repository files, materials, documentation, or the Data Availability Statement. Files supporting the published version must not be replaced silently. The journal may require a repository update, correction, reanalysis, replacement supplement, Expression of Concern, retraction, or another appropriate action.

A failure to reproduce a finding does not by itself establish misconduct. JSOMER will consider differences in methods, software, versions, stochastic variation, platforms, and access conditions. Credible concerns may lead to author requests, independent technical review, institutional contact, correction, restricted access to unsafe data, Expression of Concern, or retraction when findings are unreliable or essential evidence cannot be supplied without a credible explanation.

20Responsibilities, Compliance, and Declarations View details +

Authors are responsible for accurate records, provenance, privacy, lawful access, a complete Data Availability Statement, reasonable supporting evidence, correction of errors, third-party rights, clear distinction between observed and synthetic data, transparent reporting, and cooperation with inquiries. Reviewers must protect confidentiality, avoid unauthorized reuse or participant identification, and must not demand public release when sharing would be unethical or unlawful.

Editors apply the policy consistently, evaluate legitimate restrictions, obtain specialist advice, protect confidential data, respond to credible integrity concerns, distinguish reproducibility limitations from misconduct, and ensure that published statements are not misleading. The publisher supports publication of statements and repository links, supplementary-material preservation, metadata updates, confidentiality, and correction mechanisms but does not acquire ownership of authors’ datasets.

A manuscript may be returned, suspended, rejected, or withdrawn when a required statement is absent or misleading, legitimate restrictions are not explained, essential evidence is unavailable without credible reason, ethical or legal access terms were violated, data were fabricated or falsified, repository files do not correspond to the manuscript, or authors refuse reasonable editorial requests. Absence of open data alone is not a basis for rejection when restrictions are legitimate, transparent, and appropriately managed.

Data and Reproducibility Declaration

The authors confirm that the manuscript accurately describes the data, materials, methods, and analyses supporting its findings. The Data Availability Statement is complete and accurate. Data, code, and research materials have been shared where ethically, legally, contractually, and technically possible. Any restrictions have been explained transparently. The authors agree to provide reasonable supporting documentation upon editorial request and to notify the Journal of Social Media Research promptly if errors or availability changes are identified.

21Policy Review View details +

JSOMER reviews this policy periodically to ensure consistency with its actual editorial practices, recognized data-management and reproducibility standards, FAIR principles, transparency frameworks, relevant clinical data-sharing expectations, ethics and consent requirements, artificial-intelligence guidance, article and submission requirements, correction and retraction procedures, applicable law, institutional requirements, and relevant scholarly services.

Material changes will be disclosed through the journal website. The policy primarily requires accurate disclosure and appropriate supporting evidence, while encouraging or requiring sharing when ethically, legally, contractually, and technically possible.

Contact

Questions concerning data availability, repositories, controlled access, social media data, analytical code, reproducibility, or suspected data-integrity concerns may be sent to the JSOMER Editorial Office at editor@jsomer.org.

Last updated: July 20, 2026