ovirt questions - Page 1

Paul

Asked: 2021-05-05 09:28:27 +0800 CST

How long does it take to failover with oVirt/RHEV?

1

I'm trying to get a better understanding of what is achieved by using RHEV/oVirt (or other OSS solutions) in an HA cluster. I'm interested in knowing how long it takes to fail over, and what exactly is happening when that happens, so I can judge whether or not this is an acceptable solution for different types of situations.

For example, what's the state of the system when it comes back - is it exactly where it left off, or is it like the power was pulled from the system, and it would be restarting after a power outage (thus having inconsistent disk states?)

I know this is a bit of a vague ask... but are there best practices for VMs to be run in an HA configuration like this, with the above considerations? From a layperson coming in with little to no experience, it seems like any application should be able to just be put on a VM and it'll magically work if the primary VM host crashes, and another VM host will take over. But it seems that's not really the case, and maybe there's some fundamental considerations that can be applied to most solutions.

sloweriang

Asked: 2020-12-03 05:17:48 +0800 CST

How to configure multiple interface for cloud_init_nics using variables in ansible

1

i need some help on configure multiple cloud_init_nics using variable files. Here is my variable files for example: files/dict

vm:
  all:
    - name: rhel7
      hostname: rhel7
      dns: "8.8.8.8 8.8.4.4"
      nic:
        - nic_name: eth0
          ip: "10.10.10.10"
          netmask: "255.255.255.0"
          gateway: "10.10.10.1"
          bootproto: "static"
          onboot: "true"
        - nic_name: "eth1"
          ip: "10.10.10.11"
          netmask: "255.255.255.0"
          bootproto: "static"
          onboot: "true"

    - name: rhel8
      hostname: rhel8
      dns: "8.8.8.8 8.8.4.4"
      nic:
        - ip: "10.10.10.12"
          netmask: "255.255.255.0"
          gateway: "10.10.10.1"
          nic_name: "ens7"
          bootproto: "static"
          onboot: "true"

Each vm could have from minimum 1 nic to N number of nic card.

Here is my playbook(obviously didn't work), because it will loop the second nic as a different tasks.

- name: include vards
  include_vars: files/dict

- name: readvar
  ovirt_vm:
    cloud_init_nics:
    - nic_name: "{{item.nic|json_query('[*].nic_name') }}"
      nic_boot_protocol: "{{item.nic|json_query('[*].ip') }}"
    cloud_init_persist: no
    wait: false
  with_items: "{{vm.all}}"

From the documentation , it seems like it need multiple "- nic_name" under the same task

  cloud_init_nics:
    - nic_name: eth0
      nic_boot_protocol: dhcp
    - nic_name: eth1
      nic_boot_protocol: static
      nic_ip_address: 10.34.60.86
      nic_netmask: 255.255.252.0

Because each vm has different number of nic , so I cant use the example from the documentation.

So my question would be: how do i loop the cloud_init_nics multiple times but still running as 1 task? Is that possible? if not is there any idea that i should look up to?

is it possible to register item.all.nic as a variable, then cloud_init_nics: {{ var }} ? Thanks

Randall

Asked: 2020-03-26 10:08:03 +0800 CST

How should I recover from Ovirt and GlusterFS partial failure?

2

I am managing a 3 node Ovirt 4.3.7 cluster with a hosted engine appliance; the nodes are also glusterfs nodes. The systems are:

ovirt1 (node at 192.168.40.193)
ovirt2 (node at 192.168.40.194)
ovirt3 (node at 192.168.40.195)
ovirt-engine (engine at 192.168.40.196)

The services ovirt-ha-agent and ovirt-ha-broker are continually restarting on ovirt1 and ovirt3, and this does not seem healthy (the first notice we had of this problem were the logs for these services filling on these systems).

All indications from the GUI consoles are that overt-engine is running on ovirt3. I tried migrating overt-engine to ovirt2, but got a failure without further explanation.

Users are able to create, start, and stop VMs on all three nodes without issue.

I am seeing the following output from gluster-eventaapi status and hosted-engine --vm-status on each of the nodes:

ovirt1:

[root@ovirt1 ~]# gluster-eventsapi status
Webhooks: 
http://ovirt-engine.low.mdds.tcs-sec.com:80/ovirt-engine/services/glusterevents

+---------------+-------------+-----------------------+
|      NODE     | NODE STATUS | GLUSTEREVENTSD STATUS |
+---------------+-------------+-----------------------+
| 192.168.5.194 |          UP |                    OK |
| 192.168.5.195 |          UP |                    OK |
|   localhost   |          UP |                    OK |
+---------------+-------------+-----------------------+
[root@ovirt1 ~]# hosted-engine --vm-status
The hosted engine configuration has not been retrieved from shared storage. Please ensure that ovirt-ha-agent is running and the storage server is reachable.

ovirt2:

[root@ovirt2 ~]# gluster-eventsapi status
Webhooks: 
http://ovirt-engine.low.mdds.tcs-sec.com:80/ovirt-engine/services/glusterevents

+---------------+-------------+-----------------------+
|      NODE     | NODE STATUS | GLUSTEREVENTSD STATUS |
+---------------+-------------+-----------------------+
| 192.168.5.195 |          UP |                    OK |
| 192.168.5.193 |          UP |                    OK |
|   localhost   |          UP |                    OK |
+---------------+-------------+-----------------------+
[root@ovirt2 ~]# hosted-engine --vm-status


--== Host ovirt2.low.mdds.tcs-sec.com (id: 1) status ==--

conf_on_shared_storage             : True
Status up-to-date                  : True
Hostname                           : ovirt2.low.mdds.tcs-sec.com
Host ID                            : 1
Engine status                      : {"reason": "vm not running on this host", "health": "bad", "vm": "down_unexpected", "detail": "unknown"}
Score                              : 0
stopped                            : False
Local maintenance                  : False
crc32                              : e564d06b
local_conf_timestamp               : 9753700
Host timestamp                     : 9753700
Extra metadata (valid at timestamp):
    metadata_parse_version=1
    metadata_feature_version=1
    timestamp=9753700 (Wed Mar 25 17:45:50 2020)
    host-id=1
    score=0
    vm_conf_refresh_time=9753700 (Wed Mar 25 17:45:50 2020)
    conf_on_shared_storage=True
    maintenance=False
    state=EngineUnexpectedlyDown
    stopped=False
    timeout=Thu Apr 23 21:29:10 1970


--== Host ovirt3.low.mdds.tcs-sec.com (id: 3) status ==--

conf_on_shared_storage             : True
Status up-to-date                  : False
Hostname                           : ovirt3.low.mdds.tcs-sec.com
Host ID                            : 3
Engine status                      : unknown stale-data
Score                              : 3400
stopped                            : False
Local maintenance                  : False
crc32                              : 620c8566
local_conf_timestamp               : 1208310
Host timestamp                     : 1208310
Extra metadata (valid at timestamp):
    metadata_parse_version=1
    metadata_feature_version=1
    timestamp=1208310 (Mon Dec 16 21:14:24 2019)
    host-id=3
    score=3400
    vm_conf_refresh_time=1208310 (Mon Dec 16 21:14:24 2019)
    conf_on_shared_storage=True
    maintenance=False
    state=GlobalMaintenance
    stopped=False

ovirt3:

[root@ovirt3 ~]# gluster-eventsapi status
Webhooks: 
http://ovirt-engine.low.mdds.tcs-sec.com:80/ovirt-engine/services/glusterevents

+---------------+-------------+-----------------------+
|      NODE     | NODE STATUS | GLUSTEREVENTSD STATUS |
+---------------+-------------+-----------------------+
| 192.168.5.193 |        DOWN |           NOT OK: N/A |
| 192.168.5.194 |          UP |                    OK |
|   localhost   |          UP |                    OK |
+---------------+-------------+-----------------------+
[root@ovirt3 ~]# hosted-engine --vm-status
The hosted engine configuration has not been retrieved from shared storage. Please ensure that ovirt-ha-agent is running and the storage server is reachable.

The steps I've taken so far are:

find that logs for the ovirt-ha-agent and ovirt-ha-broker service are not rotating correctly on nodes ovirt1 and ovirt3; the logs show the same failure on both nodes. The broker.log contains this statement repeated frequently:

MainThread::WARNING::2020-03-25 18:03:28,846::storage_broker::97::ovirt_hosted_engine_ha.broker.storage_broker.StorageBroker::(__init__) Can't connect vdsm storage: [Errno 5] Input/output error: '/rhev/data-center/mnt/glusterSD/ovirt2:_engine/182a4a94-743f-4941-89c1-dc2008ae1cf5/ha_agent/hosted-engine.lockspace'

find that the RHEV documentation suggests running hosted-engine --vm-status to understand the problem; that output (above) suggests that ovirt1 is not completely part of the cluster.
I asked on the Ovirt forum yesterday morning, but since I am new there, my question needs a moderator review, and that hasn't happened yet (if the users of this cluster weren't all suddenly working from home, and suddenly dependent upon it, I wouldn't be worried about waiting a few days).

How should I recover from this situation? (I think I need to recover something in the glusterfs cluster first, but can't find a hint or don't have the language to form the right query.)

UPDATE: After restarting glusterd on ovirt3, the glusterfs cluster appears to be healthy, but with no change in behavior on the ovirt services.

itsafire

Asked: 2019-05-17 09:13:55 +0800 CST

How do I upgrade from oVirt 4.2 to oVirt 4.3?

2

I have a oVirt 4.2 datacenter. I like to upgrade to 4.3. What steps are necessary and where do I get documentation. The oVirt documentation is sparse on this matter to say the least.

Sandra

Asked: 2012-05-04 13:21:31 +0800 CST

What is the difference between RHEV and oVirt?

9

When I read the wikipedia articles about RHEV and oVirt, I can't really figure out why Red Hat have both projects, as they seam to solve the same problem?

http://en.wikipedia.org/wiki/RHEV

http://en.wikipedia.org/wiki/OVirt

oVirt will be included in Fedora 17, so clearly they have invested a lot in both projects.

Is one of them a short term solution, or do they solve different problems/tasks?

Update

Based on this link, oVirt is only a manangement interface build up from common open source packages.

So I would speculate that the management part i RHEV will be replaced with oVirt at some point.

How long does it take to failover with oVirt/RHEV?

How to configure multiple interface for cloud_init_nics using variables in ansible

How should I recover from Ovirt and GlusterFS partial failure?

How do I upgrade from oVirt 4.2 to oVirt 4.3?

What is the difference between RHEV and oVirt?

Can you pass user/pass for HTTP Basic Authentication in URL parameters?

Ping a Specific Port

Check if port is open or closed on a Linux server?

How to automate SSH login with password?

How do I tell Git for Windows where to find my private RSA key?

What's the default superuser username/password for postgres after a new install?

What port does SFTP use?

Command line to list users in a Windows Active Directory group?

What is a Pem file and how does it differ from other OpenSSL Generated Key File Formats?

How to determine if a bash variable is empty?

Questions[ovirt](server)